Bayesian latent hierarchical Imputation of Missing Ordinal Survey Responses via INLA
Thursday, Aug 6: 9:50 AM - 10:05 AM
2015
Contributed Papers
Thomas M. Menino Convention & Exhibition Center
National research platforms such as the NIH All of Us Research Program link electronic health records with participant-reported surveys, enabling large-scale real-world studies. A persistent barrier is extensive missingness in key socioeconomic survey items-especially ordinal or categorical variables such as household income-driven by nonresponse and differential participation across subpopulations and geography. Complete-case analyses and standard FCS/MICE workflows can yield biased effect estimates, poor uncertainty quantification, and limited scalability when missingness is high and heterogeneity is substantial. We propose a scalable Bayesian imputation framework that models income as a latent continuous variable mapped to observed categories through an ordered-logit measurement model, with hierarchical structure to borrow strength across geographic units (e.g., 3-digit ZIP) and incorporate area-level auxiliary information. The proposed method builds whole Bayesian model framework combines latent income model, missingness indicator model and downstream outcome model into one joint posterior problem. Performance measurements from several aspects (1) accuracy of imputed levels (2) downstream household income-outcome association recovery show that INLA captures association between 3-digit ZIP structure and latent household income/missingness, and address the uncertainty of imputation better than candidate methods like MissForest.
Missing data
Ordinal/categorical imputation
Bayesian hierarchical model with latent structure
Zip code aggregated summary statistics
the All of Us
Survey nonresponse
Main Sponsor
Section on Bayesian Statistical Science
You have unsaved changes.