Monday, Aug 3: 2:00 PM - 3:50 PM
6216
Contributed Papers
Thomas M. Menino Convention & Exhibition Center
Room: CC-209
Main Sponsor
Section on Bayesian Statistical Science
Presentations
In many scientific and public health studies, response data often comprise heterogeneous outcomes, such as continuous and binary variables, along with a set of predictors collected from experiments. It is important to account for the inherent association between these heterogeneous outcomes when simultaneously addressing parameter estimation and variable selection in scenarios where the goal is to jointly model both responses within a unified modeling framework. In this paper, we present a general strategy that employs a normal linear regression for the continuous outcome and a latent variable structure to approximate logistic regression for the binary outcome. A hybrid Gibbs sampling algorithm is developed, incorporating Metropolis–Hastings steps to efficiently update model parameters, particularly those with intractable conditional distributions. The proposed sampling strategy not only improves efficiency and reliability in statistical inference but is also well-suited for high-dimensional Bayesian models and complex datasets. Numerical simulations demonstrate the effectiveness of the proposed approach, and a real-world application is provided to illustrate its practical utility.
Keywords
Bayesian Method
Joint Hierarchical Models
Mixed Responses
Variable Selection
MCMC Sampling
We propose a computationally scalable Bayesian variable selection
framework that incorporates dependence structures among predictors
while remaining robust to hyperparameter specification. By introducing
a latent continuous variable and a Gaussian Markov random field prior
within a discrete spike-and-slab framework, our approach circumvents
the phase transition problems that typically plague binary Markov
random field priors. To ensure scalability, we develop an algorithmic
design that exploits the model's algebraic structure. This enables
fast, numerically stable updates of posterior quantities and
eliminates the need for repeated matrix inversions without
compromising inferential accuracy. We demonstrate the method's
superior performance and computational efficiency through simulation
studies and illustrate its practical utility using high-dimensional
genomic data.
Keywords
Bayesian variable selection
Spike-and-slab priors
Bayesian inference
Scalable algorithms
Predictor dependence
Large-scale genomic data
We consider the variable selection problem for linear models under the "spike and slab" priors. We consider the case where n is very large and the data-generating process cannot be described by any of the models. In this setting, we show that under mild regularity conditions, variable selection is afflicted by "Model Superinduction", a phenomenon where the posterior odds ratio favor more complicated models at a rate that is exponential in n, creating severe computational bottlenecks and selecting overly complex and uninterpretable models. We show that while this phenomenon afflicts many popular choices for the spike and slab, the choice of Student distributions is robust to Superinduction. Large sample sizes also induce additional computational costs in popular stochastic variable search approaches that renders them intractable. We propose a stochastic model search utilizing a surrogate for the posterior odds ratio under the continuous spike and slab priors that enables ultra-fast model comparison. We show that these surrogates are strongly-consistent, and when used in conjunction with Student prior distributions, result in fast and interpretable variable selection.
Keywords
Variable Selection
Spike and Slab Priors
Misspecified Models
M-open Model Comparison
Stochastic Variable Search
Bayesian Linear Models
We propose a Bayesian spatial variable selection method for massive datasets using spatially varying coefficient models. Coefficients share low-rank basis functions for dimension reduction and scalable computation. A spike and slab group lasso, SSGL, prior induces structured sparsity, selecting predictors and their spatial effects without manual tuning. We build a conjugate Bayesian model with an efficient MCMC sampler and closed-form updates. Cross-validation chooses the number of basis functions and shrinkage settings to balance fit and speed. Simulations across varied sample sizes and numbers of predictors show gains over generalized additive models and multiscale geographically weighted regression in prediction, surface recovery, and variable selection. The method lowers MSE for active coefficients, improves detection of zero effects in high dimensions, and delivers calibrated uncertainty with near nominal coverage. The approach is rigorous and scalable, with possible future extensions to non-Gaussian outcomes, spatiotemporal settings, and adaptive bases.
Keywords
Spatially Varying Coefficients
Bayesian Variable Selection
Spike-and-Slab Prior
Group Lasso
Basis Functions
Compositional predictors, such as microbiome abundances, pose unique challenges in variable selection due to their unit-sum constraint and inherent dependencies. Existing approaches often rely on fixed association graphs derived from phylogenetic or ecological distances, which may not reflect outcome-relevant relationships. We propose GRACE (GRaph-Adaptive horseshoe for Compositional rEgression), a fully Bayesian framework that enforces compositional constraints, performs variable selection, and adaptively learns an outcome-driven feature graph. GRACE achieves compositionality through a novel linear reparameterization of regression coefficients, while a structured horseshoe prior induces sparsity and smooths coefficients along the learned graph. Through extensive simulations, GRACE demonstrates competitive predictive accuracy and improved graph recovery compared with existing methods, particularly under graph misspecification. Application to oral microbiome data from the ORIGINS study identifies taxa associated with insulin resistance and reveals an outcome-driven microbial network that differs substantially from phylogenetic or co-occurrence networks.
Keywords
Compositional regression
Bayesian variable selection
Shrinkage priors
Microbiome data analysis.
Plasma proteomics offers a noninvasive and biologically informative approach for studying cognitive function in Alzheimer's disease (AD). However, statistical analysis of plasma proteomic data is challenged by extreme high dimensionality, where conventional methods using uniform shrinkage often suffer from systematic bias: they tend to over-shrink small informative signals while under-penalizing large redundant groups. To address this, we propose BSGSSS-HS, a Bayesian bi-level selection framework. By integrating spike-and-slab and horseshoe priors, it performs simultaneous group and within-group selection, achieving group-aware, size-invariant shrinkage under heterogeneous sparsity. For posterior inference, we employ an efficient Gibbs sampler. Extensive simulation studies show that the proposed method consistently outperforms competing approaches across a variety of scenarios, including varying group sizes, signal-to-noise ratios, and dimensionalities. We further apply BSGSSS-HS to real AD data from the Vanderbilt Memory and Aging Project cohort, where the method is able to identify proteins and pathways implicated in AD-related pathological processes and cognitive decline.
Keywords
High-dimensional
Bi-level selection
Bayesian variable selection
Horseshoe prior
spike and slab prior
Heterogeneous sparsity