Efficient quasi-randomization of administrative data partially linked to a probability sample from the same population

Vladislav Beresovsky Speaker
U.S. Bureau of Labor Statistics
 
Julie Gershunskaya Co-Author
US Bureau of Labor Statistics
 
Terrance Savitsky Co-Author
US Bureau of Labor Statistics
 
Tuesday, Aug 4: 3:20 PM - 3:45 PM
Invited Paper Session 
Thomas M. Menino Convention & Exhibition Center 
Lowering response rates and raising costs of traditional probability-based surveys motivate increased interest in using nonprobability data sources, such as web surveys and administrative records, to produce estimates of target population quantities. Methods have been developed to account for a selection bias associated with such "convenience" nonprobability data. We consider estimation of response propensity to nonprobability dataset by combining it with a probability "reference" sample obtained from the same target population and maximizing Bernoulli likelihood for the observed sample indicators. We use the missing information principle (MIP) to utilize information from probabilistic data linkage to improve robustness and efficiency of the estimated response propensity. We compare our proposed method with a commonly used pseudo-likelihood approach.

Keywords

Design-based inference

Non-probability sample

Response propensity

Probabilistic data linkage

Missing information principle

Implicit logistic regression