Massive Samples, Misspecified Models: Robust and Tractable Spike and Slab Variable Selection

Jacob Fontana Speaker
 
Bruno Sanso Co-Author
University of California-Santa Cruz
 
Monday, Aug 3: 2:35 PM - 2:50 PM
3740 
Contributed Papers 
Thomas M. Menino Convention & Exhibition Center 
We consider the variable selection problem for linear models under the "spike and slab" priors. We consider the case where n is very large and the data-generating process cannot be described by any of the models. In this setting, we show that under mild regularity conditions, variable selection is afflicted by "Model Superinduction", a phenomenon where the posterior odds ratio favor more complicated models at a rate that is exponential in n, creating severe computational bottlenecks and selecting overly complex and uninterpretable models. We show that while this phenomenon afflicts many popular choices for the spike and slab, the choice of Student distributions is robust to Superinduction. Large sample sizes also induce additional computational costs in popular stochastic variable search approaches that renders them intractable. We propose a stochastic model search utilizing a surrogate for the posterior odds ratio under the continuous spike and slab priors that enables ultra-fast model comparison. We show that these surrogates are strongly-consistent, and when used in conjunction with Student prior distributions, result in fast and interpretable variable selection.

Keywords

Variable Selection

Spike and Slab Priors

Misspecified Models

M-open Model Comparison

Stochastic Variable Search

Bayesian Linear Models 

Main Sponsor

Section on Bayesian Statistical Science