Bayesian latent hierarchical Imputation of Missing Ordinal Survey Responses via INLA

Haoyang Yi Speaker
 
Qingxia Chen Co-Author
Vanderbilt University Medical Center
 
Thursday, Aug 6: 9:50 AM - 10:05 AM
2015 
Contributed Papers 
Thomas M. Menino Convention & Exhibition Center 
National research platforms such as the NIH All of Us Research Program link electronic health records with participant-reported surveys, enabling large-scale real-world studies. A persistent barrier is extensive missingness in key socioeconomic survey items-especially ordinal or categorical variables such as household income-driven by nonresponse and differential participation across subpopulations and geography. Complete-case analyses and standard FCS/MICE workflows can yield biased effect estimates, poor uncertainty quantification, and limited scalability when missingness is high and heterogeneity is substantial. We propose a scalable Bayesian imputation framework that models income as a latent continuous variable mapped to observed categories through an ordered-logit measurement model, with hierarchical structure to borrow strength across geographic units (e.g., 3-digit ZIP) and incorporate area-level auxiliary information. The proposed method builds whole Bayesian model framework combines latent income model, missingness indicator model and downstream outcome model into one joint posterior problem. Performance measurements from several aspects (1) accuracy of imputed levels (2) downstream household income-outcome association recovery show that INLA captures association between 3-digit ZIP structure and latent household income/missingness, and address the uncertainty of imputation better than candidate methods like MissForest.

Keywords

Missing data

Ordinal/categorical imputation

Bayesian hierarchical model with latent structure

Zip code aggregated summary statistics

the All of Us

Survey nonresponse 

Main Sponsor

Section on Bayesian Statistical Science