Maximizing continuous auxiliary data for improved survey inference while controlling disclosure risk

Sharifa Williams Speaker
 
Jungang Zou Co-Author
 
Yutao Liu Co-Author
 
Yajuan Si Co-Author
University of Michigan
 
Sandro Galea Co-Author
Washington University School of Public Health
 
Qixuan Chen Co-Author
Columbia University
 
Wednesday, Aug 5: 9:05 AM - 9:20 AM
2595 
Contributed Papers 
Thomas M. Menino Convention & Exhibition Center 
Probability surveys face rising non-response rates, resulting in biased statistical inference. Auxiliary information can be used to reduce bias in estimation. Continuous auxiliary variables in administrative data are often discretized prior to release to avoid confidentiality breaches. This may limit the utility of administrative records in improving survey estimates, especially when continuous auxiliary data strongly predict the survey outcome. We propose a two-step strategy. First, statistical agencies use confidential continuous auxiliary data to estimate response propensity score of the survey sample and include these in a modified population data. Data users then conduct predictive inference including the discretized continuous variables and the propensity scores as predictors using splines in a Bayesian model. The proposed method performs well, yielding more efficient estimates of population means with 95% credible intervals providing better coverage than alternative approaches. We illustrate the proposed method using the Ohio Army National Guard Mental Health Initiative. The methods developed in this work are readily available in the R package AuxSurvey.

Keywords

Bayesian generalized additive model

continuous auxiliary variables

data protection

Rstan

response propensity

poststratification 

Main Sponsor

Survey Research Methods Section