Best practices for leveraging nonprobability samples

Daifeng Han Chair
Westat
 
Qixuan Chen Discussant
Columbia University
 
Andreea Erciulescu Organizer
Westat
 
Tuesday, Aug 4: 2:00 PM - 3:50 PM
1177 
Invited Paper Session 
Thomas M. Menino Convention & Exhibition Center 
Room: CC-102A 

Applied

No

Main Sponsor

Survey Research Methods Section

Co Sponsors

Government Statistics Section
Social Statistics Section

Presentations

Estimation Strategies with Nonignorable Non-Probability Survey Samples

We first provide an overview on inferential frameworks for analyzing non-probability survey samples. We then present some recent results on estimating participation probabilities (i.e., propensity scores) under ignorable or nonignorable participation mechanisms. In particular, we show that the pseudo maximum likelihood method of Chen, Li and Wu (2020, JASA), which was developed under the ignorability assumption, can be used to build a method for dealing with nonignorable participation mechanisms when there is a moderate or strong correlation between the study variable and auxiliary variables. Some empirical results from simulation studies will be presented.  

Keywords

Inverse probability weighting

Doubly robust estimation

Nonignorable participation

Pseudo maximum likelihood method

Instrumental variable 

Speaker

Changbao Wu, University of Waterloo

Addressing Reporting Challenges for Private School Samples in NAEP by Enhancing Representativeness Through Combining Years

The National Assessment of Educational Progress (NAEP), often referred to as "The Nation's Report Card," provides essential insights into student achievement across the United States. However, reporting challenges persist, particularly for small subgroups and low response rates in private schools. This presentation introduces a methodology developed in collaboration with NCES to address these issues. For private schools, we propose a substitution approach to use matched respondents from prior years to mitigate nonresponse bias. Using exact matching on key school characteristics and bias assessment metrics such as Hellinger distance and Kullback-Leibler divergence, the approach demonstrates promising improvements in representativeness and meeting response rate thresholds. While effective for earlier years, limitations arise with declining response rates in recent cycles, prompting further evaluation. These strategies offer insights for sustaining the integrity and utility of NAEP reporting. 

Keywords

Nonresponse bias

Sample substitution 

Speaker

Thomas Krenzke, Westat

Estimation of a discrete distribution function from a non-probability sample: an uncertainty based approach

In the latest decade, the relevance of non-probability samples is considerably increased because
of the decay of response rates in traditional sample surveys, and the availability of massive
datasets at a relatively low cost. In probability sampling, each unit of the population has
a known, non-zero probability of being selected. Non-probability samples involve, on the contrary,
uncontrolled methods for selecting of units into the sample. This implies that inclusion probabilities
are unknown, and then it is not possible, through the Inverse Probability Weighting principle, to
remove the selection bias. Furthermore, the unknown selection mechanism is frequently selective
with respect to the target population, so estimates of population characteristics may be subject to
serious selection bias. The basic question, therefore, is how to draw inference from such samples,
for population parameters of interest.
As a major effect of the uncontrolled selection mechanism, unless very restrictive assumptions are
made, the distribution of the character of interest is unidentifiable. In this paper the concept of
uncertainty on data generating model, resulting from the lack of knowledge of the sampling design
acting in the non-probability sample, is introduced. Furthermore, the reduction of uncertainty due to
the availability of extra-sample in formation is discussed.  

Keywords

informative sample

non-probability sample

uncertainty 

Speaker

Daniela Marella, Sapienza University of Rome

Co-Author

Pier Luigi Conti, Dipart Di Stat Prob E Stat

Efficient quasi-randomization of administrative data partially linked to a probability sample from the same population

Lowering response rates and raising costs of traditional probability-based surveys motivate increased interest in using nonprobability data sources, such as web surveys and administrative records, to produce estimates of target population quantities. Methods have been developed to account for a selection bias associated with such "convenience" nonprobability data. We consider estimation of response propensity to nonprobability dataset by combining it with a probability "reference" sample obtained from the same target population and maximizing Bernoulli likelihood for the observed sample indicators. We use the missing information principle (MIP) to utilize information from probabilistic data linkage to improve robustness and efficiency of the estimated response propensity. We compare our proposed method with a commonly used pseudo-likelihood approach.  

Keywords

Design-based inference

Non-probability sample

Response propensity

Probabilistic data linkage

Missing information principle

Implicit logistic regression 

Speaker

Vladislav Beresovsky, U.S. Bureau of Labor Statistics

Co-Author(s)

Julie Gershunskaya, US Bureau of Labor Statistics
Terrance Savitsky, US Bureau of Labor Statistics