Tuesday, Aug 4: 2:00 PM - 3:50 PM
1177
Invited Paper Session
Thomas M. Menino Convention & Exhibition Center
Room: CC-102A
Applied
No
Main Sponsor
Survey Research Methods Section
Co Sponsors
Government Statistics Section
Social Statistics Section
Presentations
We first provide an overview on inferential frameworks for analyzing non-probability survey samples. We then present some recent results on estimating participation probabilities (i.e., propensity scores) under ignorable or nonignorable participation mechanisms. In particular, we show that the pseudo maximum likelihood method of Chen, Li and Wu (2020, JASA), which was developed under the ignorability assumption, can be used to build a method for dealing with nonignorable participation mechanisms when there is a moderate or strong correlation between the study variable and auxiliary variables. Some empirical results from simulation studies will be presented.
Keywords
Inverse probability weighting
Doubly robust estimation
Nonignorable participation
Pseudo maximum likelihood method
Instrumental variable
The National Assessment of Educational Progress (NAEP), often referred to as "The Nation's Report Card," provides essential insights into student achievement across the United States. However, reporting challenges persist, particularly for small subgroups and low response rates in private schools. This presentation introduces a methodology developed in collaboration with NCES to address these issues. For private schools, we propose a substitution approach to use matched respondents from prior years to mitigate nonresponse bias. Using exact matching on key school characteristics and bias assessment metrics such as Hellinger distance and Kullback-Leibler divergence, the approach demonstrates promising improvements in representativeness and meeting response rate thresholds. While effective for earlier years, limitations arise with declining response rates in recent cycles, prompting further evaluation. These strategies offer insights for sustaining the integrity and utility of NAEP reporting.
Keywords
Nonresponse bias
Sample substitution
In the latest decade, the relevance of non-probability samples is considerably increased because
of the decay of response rates in traditional sample surveys, and the availability of massive
datasets at a relatively low cost. In probability sampling, each unit of the population has
a known, non-zero probability of being selected. Non-probability samples involve, on the contrary,
uncontrolled methods for selecting of units into the sample. This implies that inclusion probabilities
are unknown, and then it is not possible, through the Inverse Probability Weighting principle, to
remove the selection bias. Furthermore, the unknown selection mechanism is frequently selective
with respect to the target population, so estimates of population characteristics may be subject to
serious selection bias. The basic question, therefore, is how to draw inference from such samples,
for population parameters of interest.
As a major effect of the uncontrolled selection mechanism, unless very restrictive assumptions are
made, the distribution of the character of interest is unidentifiable. In this paper the concept of
uncertainty on data generating model, resulting from the lack of knowledge of the sampling design
acting in the non-probability sample, is introduced. Furthermore, the reduction of uncertainty due to
the availability of extra-sample in formation is discussed.
Keywords
informative sample
non-probability sample
uncertainty
Lowering response rates and raising costs of traditional probability-based surveys motivate increased interest in using nonprobability data sources, such as web surveys and administrative records, to produce estimates of target population quantities. Methods have been developed to account for a selection bias associated with such "convenience" nonprobability data. We consider estimation of response propensity to nonprobability dataset by combining it with a probability "reference" sample obtained from the same target population and maximizing Bernoulli likelihood for the observed sample indicators. We use the missing information principle (MIP) to utilize information from probabilistic data linkage to improve robustness and efficiency of the estimated response propensity. We compare our proposed method with a commonly used pseudo-likelihood approach.
Keywords
Design-based inference
Non-probability sample
Response propensity
Probabilistic data linkage
Missing information principle
Implicit logistic regression