Bayesian Variable Selection and Shrinkage Methods

Christine Peterson Chair
Rice University
 
Monday, Aug 3: 2:00 PM - 3:50 PM
6216 
Contributed Papers 
Thomas M. Menino Convention & Exhibition Center 
Room: CC-209 

Main Sponsor

Section on Bayesian Statistical Science

Presentations

Bayesian Variable Selection Regression with Quantitative and Qualitative Outcomes

In many scientific and public health studies, response data often comprise heterogeneous outcomes, such as continuous and binary variables, along with a set of predictors collected from experiments. It is important to account for the inherent association between these heterogeneous outcomes when simultaneously addressing parameter estimation and variable selection in scenarios where the goal is to jointly model both responses within a unified modeling framework. In this paper, we present a general strategy that employs a normal linear regression for the continuous outcome and a latent variable structure to approximate logistic regression for the binary outcome. A hybrid Gibbs sampling algorithm is developed, incorporating Metropolis–Hastings steps to efficiently update model parameters, particularly those with intractable conditional distributions. The proposed sampling strategy not only improves efficiency and reliability in statistical inference but is also well-suited for high-dimensional Bayesian models and complex datasets. Numerical simulations demonstrate the effectiveness of the proposed approach, and a real-world application is provided to illustrate its practical utility. 

Keywords

Bayesian Method

Joint Hierarchical Models

Mixed Responses

Variable Selection

MCMC Sampling 

Speaker

Min Wang, University of Texas At San Antonio

Co-Author(s)

Mai Dao, Wichita State University
Prince Buti

Efficient Bayesian Variable Selection under Predictor Dependence Regimes

We propose a computationally scalable Bayesian variable selection
framework that incorporates dependence structures among predictors
while remaining robust to hyperparameter specification. By introducing
a latent continuous variable and a Gaussian Markov random field prior
within a discrete spike-and-slab framework, our approach circumvents
the phase transition problems that typically plague binary Markov
random field priors. To ensure scalability, we develop an algorithmic
design that exploits the model's algebraic structure. This enables
fast, numerically stable updates of posterior quantities and
eliminates the need for repeated matrix inversions without
compromising inferential accuracy. We demonstrate the method's
superior performance and computational efficiency through simulation
studies and illustrate its practical utility using high-dimensional
genomic data. 

Keywords

Bayesian variable selection

Spike-and-slab priors

Bayesian inference

Scalable algorithms

Predictor dependence

Large-scale genomic data 

Speaker

Farshid Abadizaman

Co-Author

Mahlet Tadesse, Georgetown University

Massive Samples, Misspecified Models: Robust and Tractable Spike and Slab Variable Selection

We consider the variable selection problem for linear models under the "spike and slab" priors. We consider the case where n is very large and the data-generating process cannot be described by any of the models. In this setting, we show that under mild regularity conditions, variable selection is afflicted by "Model Superinduction", a phenomenon where the posterior odds ratio favor more complicated models at a rate that is exponential in n, creating severe computational bottlenecks and selecting overly complex and uninterpretable models. We show that while this phenomenon afflicts many popular choices for the spike and slab, the choice of Student distributions is robust to Superinduction. Large sample sizes also induce additional computational costs in popular stochastic variable search approaches that renders them intractable. We propose a stochastic model search utilizing a surrogate for the posterior odds ratio under the continuous spike and slab priors that enables ultra-fast model comparison. We show that these surrogates are strongly-consistent, and when used in conjunction with Student prior distributions, result in fast and interpretable variable selection. 

Keywords

Variable Selection

Spike and Slab Priors

Misspecified Models

M-open Model Comparison

Stochastic Variable Search

Bayesian Linear Models 

Speaker

Jacob Fontana

Co-Author

Bruno Sanso, University of California-Santa Cruz

Bayesian Variable Selection for Spatially Varying Coefficient Models via Spike-and-Slab Group Lasso

We propose a Bayesian spatial variable selection method for massive datasets using spatially varying coefficient models. Coefficients share low-rank basis functions for dimension reduction and scalable computation. A spike and slab group lasso, SSGL, prior induces structured sparsity, selecting predictors and their spatial effects without manual tuning. We build a conjugate Bayesian model with an efficient MCMC sampler and closed-form updates. Cross-validation chooses the number of basis functions and shrinkage settings to balance fit and speed. Simulations across varied sample sizes and numbers of predictors show gains over generalized additive models and multiscale geographically weighted regression in prediction, surface recovery, and variable selection. The method lowers MSE for active coefficients, improves detection of zero effects in high dimensions, and delivers calibrated uncertainty with near nominal coverage. The approach is rigorous and scalable, with possible future extensions to non-Gaussian outcomes, spatiotemporal settings, and adaptive bases. 

Keywords

Spatially Varying Coefficients

Bayesian Variable Selection

Spike-and-Slab Prior

Group Lasso

Basis Functions 

Speaker

Cheng-Han Yu, Marquette University

Co-Author

Qishi Zhan, Marquette University

Graph-Adaptive Shrinkage for Compositional Regression

Compositional predictors, such as microbiome abundances, pose unique challenges in variable selection due to their unit-sum constraint and inherent dependencies. Existing approaches often rely on fixed association graphs derived from phylogenetic or ecological distances, which may not reflect outcome-relevant relationships. We propose GRACE (GRaph-Adaptive horseshoe for Compositional rEgression), a fully Bayesian framework that enforces compositional constraints, performs variable selection, and adaptively learns an outcome-driven feature graph. GRACE achieves compositionality through a novel linear reparameterization of regression coefficients, while a structured horseshoe prior induces sparsity and smooths coefficients along the learned graph. Through extensive simulations, GRACE demonstrates competitive predictive accuracy and improved graph recovery compared with existing methods, particularly under graph misspecification. Application to oral microbiome data from the ORIGINS study identifies taxa associated with insulin resistance and reveals an outcome-driven microbial network that differs substantially from phylogenetic or co-occurrence networks. 

Keywords

Compositional regression

Bayesian variable selection

Shrinkage priors

Microbiome data analysis. 

Speaker

Satabdi Saha, The University of Texas MD Anderson Cancer Center

Co-Author

Christine Peterson, Rice University

Novel Bayesian Bi-Level Variable Selection Method with Application to Alzheimer’s Disease

Plasma proteomics offers a noninvasive and biologically informative approach for studying cognitive function in Alzheimer's disease (AD). However, statistical analysis of plasma proteomic data is challenged by extreme high dimensionality, where conventional methods using uniform shrinkage often suffer from systematic bias: they tend to over-shrink small informative signals while under-penalizing large redundant groups. To address this, we propose BSGSSS-HS, a Bayesian bi-level selection framework. By integrating spike-and-slab and horseshoe priors, it performs simultaneous group and within-group selection, achieving group-aware, size-invariant shrinkage under heterogeneous sparsity. For posterior inference, we employ an efficient Gibbs sampler. Extensive simulation studies show that the proposed method consistently outperforms competing approaches across a variety of scenarios, including varying group sizes, signal-to-noise ratios, and dimensionalities. We further apply BSGSSS-HS to real AD data from the Vanderbilt Memory and Aging Project cohort, where the method is able to identify proteins and pathways implicated in AD-related pathological processes and cognitive decline. 

Keywords

High-dimensional

Bi-level selection

Bayesian variable selection

Horseshoe prior

spike and slab prior

Heterogeneous sparsity 

Speaker

Manjun Yu

Co-Author(s)

Xiaojing Wang, University of Connecticut
Panpan Zhang, Vanderbilt University Medical Center