Bayesian Methods for Categorical and Survey Data

Abhi Jain Chair
 
Thursday, Aug 6: 8:30 AM - 10:20 AM
6221 
Contributed Papers 
Thomas M. Menino Convention & Exhibition Center 
Room: CC-254B 

Main Sponsor

Section on Bayesian Statistical Science

Presentations

Valid and Efficient Possibilistic Inferential Models for Categorical Data

In this exposé of the inferential model (IM), we investigate how finite-sample calibrated inference can be done when you have categorical data. IMs produce plausibility and necessity measures that are provably reliable for all sample sizes while simultaneously providing a Bayesian-like interpretable output. For multinomial and categorical regression inference, we show how plausibility contours derived from validified relative likelihoods yield regions with guaranteed frequentist coverage.

A key challenge in the applications of IMs to categorical data problems is computation. Discrete models produce "jumps" in plausibility that invalidate existing methods aimed at doing gradient descent on the contour to find a probabilistic approximation to the IM. We propose using an importance sampling procedure to amortize plausibility evaluations. We provide guidance on using the importance sampler in the multinomial and categorical regression problems using questions about odds ratio and the Iris dataset as motivating examples. 

Keywords

categorical data

Bayesian

possibility theory

importance sampling

multinomial

categorical regression 

Speaker

Sahil Patel

Co-Author

Ryan Martin

Bayesian Quantile Regression for Misclassified Binary Data

Quantile regression provides a flexible framework for characterizing covariate effects across the entire conditional distribution of outcomes, rather than focusing solely on the mean. In studies with binary outcomes, the observed response is often subject to misclassification, which can introduce substantial bias. To address this challenge, we propose a Bayesian quantile regression framework that explicitly accounts for misclassification in binary outcomes by introducing a latent true outcome and incorporating sensitivity and specificity parameters. Extensive simulation studies are conducted under varying degrees of misclassification, prior specifications, and effective sample sizes. The proposed approach is illustrated through an application examining the effect of female employment status on the likelihood of domestic violence. 

Keywords

Quantile regression

Misclassification

Markov chain Monte Carlo

Bayesian methods 

Speaker

Joon Jin Song, Baylor University

Co-Author(s)

Arshad Rahman, Indian Institute of Technology Kanpur
Yoo-Mi Chin, Baylor University
James Stamey, Baylor University

Bayesian Gaussian Copula Graphical Models for Ordinal Survey Data

It is challenging to assess conditional dependence among a large set of discrete random variables. Bayesian Gaussian copula graphical model is applied to estimate the conditional dependence for ordinal variables. Following the idea of the graphical Lasso prior, graphical spike-and-slab Lasso prior is proposed for the regularization purpose. A block Gibbs sampling scheme is then developed for the posterior computation. We further extend the graphical spike-and-slab Lasso prior to an adaptive version by allowing the parameter of the spike component tying to a prior.
Simulation study is conducted to compare the performance of using different priors.
Simulation results show a good estimation performance of estimating the latent precision matrix when using the graphical Lasso prior and the graphical spike-slab Lasso prior with a relatively large spike parameter. In terms of the graph structure learning, the adaptive graphical Lasso prior and the adaptive graphical spike-slab Lasso prior have a better performance. We then utilize the proposed methods to analyze the survey data about the physiological and psychological health of current Chinese college students. 

Keywords

Conditional independence

Gaussian copula

Graphical Lasso

Partial correlation

Precision matrix

Spike-and-slab 

Speaker

Xiaoyan Lin, University of South Carolina

Co-Author

Yang He, Department of Statistics, University of South Carolina

A Bayesian INLA Framework for Modeling Dependence in Best-Worst Discrete Choice Experiments

Analyzing highly dependent best-worst (BW) choice pairs in discrete choice experiments presents a significant challenge in complex, context-dependent settings, particularly when comparing alternative strategies and latent utilities over time. We develop a Bayesian framework for modeling dependent BW choice data in which outcomes are represented as directed transitions between latent preference states. Transition counts are modeled using a Poisson log-linear specification with pair-specific random effects to capture heterogeneity and cross-alternative dependence. Bayesian inference is conducted using Integrated Nested Laplace Approximation (INLA), enabling scalable and computationally efficient estimation without reliance on Markov chain Monte Carlo. The approach is evaluated through simulation studies and a quality-of-life case study inspired by Flynn et al., and is bench-marked against copula-based transition models. Results demonstrate improved model fit and stability, as assessed by DIC and WAIC, while providing robust and interpretable estimates of latent utilities across a range of sample sizes. 

Keywords

Best-Worst Scaling

Discrete Choice Experiments

Bayesian INLA

Utility Modeling 

Speaker

Nadeesha Jayaweera, University of Akron

Co-Author(s)

Sasanka Adikari, AOPC
Jian Zou, University of Central Florida
Norou Diawara, Old Dominion University

Bayesian Regression Analysis of Bivariate Group-Tested Current Status Data with Spatial Effects

Group testing has been a cost-effective strategy for screening rare diseases on large-scale populations. The strategy works by combining the specimens of blood or urine from multiple subjects and testing the pooled specimens instead of individual's. When a disease onset time is of interest, such group testing study design produces group-tested current status data, in which neither individual disease time nor the individual disease status is available, and only the group disease status is available. Our project studies joint analysis of two disease onset times simultaneously, for each of which only group-tested current status data are available. A new frailty model is proposed to incorporate both the dependence between correlated disease onset times and the spatial dependence among subjects sharing the same clinic. A fully Bayesian estimation approach is developed based on a data augmentation that lead to a complete data likelihood in an appealing form. The proposed Gibbs sampler is computationally efficient since all the latent variables and parameters are sampled from some well recognized full conditional distributions. Our method is evaluated by a simulation study and illustrate 

Keywords

Bivariate group-tested current status data

Frailty model

Misspecification

Gibbs sampler 

Speaker

Shuqi Song

Co-Author(s)

Lianming Wang, University of South Carolina
Christopher McMahan
Joshua Tebbs, University of South Carolina
Jihyun Kim, University of Minnesota

Bayesian latent hierarchical Imputation of Missing Ordinal Survey Responses via INLA

National research platforms such as the NIH All of Us Research Program link electronic health records with participant-reported surveys, enabling large-scale real-world studies. A persistent barrier is extensive missingness in key socioeconomic survey items-especially ordinal or categorical variables such as household income-driven by nonresponse and differential participation across subpopulations and geography. Complete-case analyses and standard FCS/MICE workflows can yield biased effect estimates, poor uncertainty quantification, and limited scalability when missingness is high and heterogeneity is substantial. We propose a scalable Bayesian imputation framework that models income as a latent continuous variable mapped to observed categories through an ordered-logit measurement model, with hierarchical structure to borrow strength across geographic units (e.g., 3-digit ZIP) and incorporate area-level auxiliary information. The proposed method builds whole Bayesian model framework combines latent income model, missingness indicator model and downstream outcome model into one joint posterior problem. Performance measurements from several aspects (1) accuracy of imputed levels (2) downstream household income-outcome association recovery show that INLA captures association between 3-digit ZIP structure and latent household income/missingness, and address the uncertainty of imputation better than candidate methods like MissForest. 

Keywords

Missing data

Ordinal/categorical imputation

Bayesian hierarchical model with latent structure

Zip code aggregated summary statistics

the All of Us

Survey nonresponse 

Speaker

Haoyang Yi

Co-Author

Qingxia Chen, Vanderbilt University Medical Center

A Bayesian Empirical Likelihood Approach to Complex Survey and Non-Probability Data

This paper proposes a Bayesian empirical likelihood (BEL) framework for analyzing complex survey data and extends it to non-probability sampling. Standard parametric likelihood methods are often hard to use with complex survey designs because the likelihood is rarely available in closed form. Empirical likelihood offers a flexible alternative by replacing the parametric likelihood with likelihoods based on moment conditions.
The proposed approach incorporates survey design features directly into the empirical likelihood framework. It is then extended to non-probability samples using selection models and design-consistent constraints. Posterior inference is conducted using a Metropolis–Hastings MCMC algorithm. A real-data application demonstrates that the proposed method can reduce selection bias when combining probability and non-probability samples. 

Keywords

Bayesian empirical likelihood

Complex survey designs

non-probability sampling

MCMC algorithm

Metropolis–Hastings MCMC 

Speaker

Md Hasibur Rahman

Co-Author

Sanjay Chaudhuri, University of Nebraska-Lincoln