Thursday, Aug 6: 8:30 AM - 10:20 AM
6221
Contributed Papers
Thomas M. Menino Convention & Exhibition Center
Room: CC-254B
Main Sponsor
Section on Bayesian Statistical Science
Presentations
In this exposé of the inferential model (IM), we investigate how finite-sample calibrated inference can be done when you have categorical data. IMs produce plausibility and necessity measures that are provably reliable for all sample sizes while simultaneously providing a Bayesian-like interpretable output. For multinomial and categorical regression inference, we show how plausibility contours derived from validified relative likelihoods yield regions with guaranteed frequentist coverage.
A key challenge in the applications of IMs to categorical data problems is computation. Discrete models produce "jumps" in plausibility that invalidate existing methods aimed at doing gradient descent on the contour to find a probabilistic approximation to the IM. We propose using an importance sampling procedure to amortize plausibility evaluations. We provide guidance on using the importance sampler in the multinomial and categorical regression problems using questions about odds ratio and the Iris dataset as motivating examples.
Keywords
categorical data
Bayesian
possibility theory
importance sampling
multinomial
categorical regression
Quantile regression provides a flexible framework for characterizing covariate effects across the entire conditional distribution of outcomes, rather than focusing solely on the mean. In studies with binary outcomes, the observed response is often subject to misclassification, which can introduce substantial bias. To address this challenge, we propose a Bayesian quantile regression framework that explicitly accounts for misclassification in binary outcomes by introducing a latent true outcome and incorporating sensitivity and specificity parameters. Extensive simulation studies are conducted under varying degrees of misclassification, prior specifications, and effective sample sizes. The proposed approach is illustrated through an application examining the effect of female employment status on the likelihood of domestic violence.
Keywords
Quantile regression
Misclassification
Markov chain Monte Carlo
Bayesian methods
It is challenging to assess conditional dependence among a large set of discrete random variables. Bayesian Gaussian copula graphical model is applied to estimate the conditional dependence for ordinal variables. Following the idea of the graphical Lasso prior, graphical spike-and-slab Lasso prior is proposed for the regularization purpose. A block Gibbs sampling scheme is then developed for the posterior computation. We further extend the graphical spike-and-slab Lasso prior to an adaptive version by allowing the parameter of the spike component tying to a prior.
Simulation study is conducted to compare the performance of using different priors.
Simulation results show a good estimation performance of estimating the latent precision matrix when using the graphical Lasso prior and the graphical spike-slab Lasso prior with a relatively large spike parameter. In terms of the graph structure learning, the adaptive graphical Lasso prior and the adaptive graphical spike-slab Lasso prior have a better performance. We then utilize the proposed methods to analyze the survey data about the physiological and psychological health of current Chinese college students.
Keywords
Conditional independence
Gaussian copula
Graphical Lasso
Partial correlation
Precision matrix
Spike-and-slab
Speaker
Xiaoyan Lin, University of South Carolina
Co-Author
Yang He, Department of Statistics, University of South Carolina
Analyzing highly dependent best-worst (BW) choice pairs in discrete choice experiments presents a significant challenge in complex, context-dependent settings, particularly when comparing alternative strategies and latent utilities over time. We develop a Bayesian framework for modeling dependent BW choice data in which outcomes are represented as directed transitions between latent preference states. Transition counts are modeled using a Poisson log-linear specification with pair-specific random effects to capture heterogeneity and cross-alternative dependence. Bayesian inference is conducted using Integrated Nested Laplace Approximation (INLA), enabling scalable and computationally efficient estimation without reliance on Markov chain Monte Carlo. The approach is evaluated through simulation studies and a quality-of-life case study inspired by Flynn et al., and is bench-marked against copula-based transition models. Results demonstrate improved model fit and stability, as assessed by DIC and WAIC, while providing robust and interpretable estimates of latent utilities across a range of sample sizes.
Keywords
Best-Worst Scaling
Discrete Choice Experiments
Bayesian INLA
Utility Modeling
Group testing has been a cost-effective strategy for screening rare diseases on large-scale populations. The strategy works by combining the specimens of blood or urine from multiple subjects and testing the pooled specimens instead of individual's. When a disease onset time is of interest, such group testing study design produces group-tested current status data, in which neither individual disease time nor the individual disease status is available, and only the group disease status is available. Our project studies joint analysis of two disease onset times simultaneously, for each of which only group-tested current status data are available. A new frailty model is proposed to incorporate both the dependence between correlated disease onset times and the spatial dependence among subjects sharing the same clinic. A fully Bayesian estimation approach is developed based on a data augmentation that lead to a complete data likelihood in an appealing form. The proposed Gibbs sampler is computationally efficient since all the latent variables and parameters are sampled from some well recognized full conditional distributions. Our method is evaluated by a simulation study and illustrate
Keywords
Bivariate group-tested current status data
Frailty model
Misspecification
Gibbs sampler
National research platforms such as the NIH All of Us Research Program link electronic health records with participant-reported surveys, enabling large-scale real-world studies. A persistent barrier is extensive missingness in key socioeconomic survey items-especially ordinal or categorical variables such as household income-driven by nonresponse and differential participation across subpopulations and geography. Complete-case analyses and standard FCS/MICE workflows can yield biased effect estimates, poor uncertainty quantification, and limited scalability when missingness is high and heterogeneity is substantial. We propose a scalable Bayesian imputation framework that models income as a latent continuous variable mapped to observed categories through an ordered-logit measurement model, with hierarchical structure to borrow strength across geographic units (e.g., 3-digit ZIP) and incorporate area-level auxiliary information. The proposed method builds whole Bayesian model framework combines latent income model, missingness indicator model and downstream outcome model into one joint posterior problem. Performance measurements from several aspects (1) accuracy of imputed levels (2) downstream household income-outcome association recovery show that INLA captures association between 3-digit ZIP structure and latent household income/missingness, and address the uncertainty of imputation better than candidate methods like MissForest.
Keywords
Missing data
Ordinal/categorical imputation
Bayesian hierarchical model with latent structure
Zip code aggregated summary statistics
the All of Us
Survey nonresponse
This paper proposes a Bayesian empirical likelihood (BEL) framework for analyzing complex survey data and extends it to non-probability sampling. Standard parametric likelihood methods are often hard to use with complex survey designs because the likelihood is rarely available in closed form. Empirical likelihood offers a flexible alternative by replacing the parametric likelihood with likelihoods based on moment conditions.
The proposed approach incorporates survey design features directly into the empirical likelihood framework. It is then extended to non-probability samples using selection models and design-consistent constraints. Posterior inference is conducted using a Metropolis–Hastings MCMC algorithm. A real-data application demonstrates that the proposed method can reduce selection bias when combining probability and non-probability samples.
Keywords
Bayesian empirical likelihood
Complex survey designs
non-probability sampling
MCMC algorithm
Metropolis–Hastings MCMC