SRMS/SSS/GSS Student Paper Competition Winners

Elizabeth Carson Chair
Science & Technology Directorate, U.S. Department of Homeland Security
 
Elizabeth Carson Organizer
Science & Technology Directorate, U.S. Department of Homeland Security
 
James Wagner Organizer
University of Michigan
 
Eli Ben-Michael Organizer
Carnegie Mellon University
 
Wednesday, Aug 5: 10:30 AM - 12:20 PM
1515 
Topic-Contributed Paper Session 
Thomas M. Menino Convention & Exhibition Center 
Room: CC-258C 
The authors of the five winning papers of the SRMS/SSS/GSS Student Paper Competition to present their research at this session. The papers were selected from over 30 submitted to SRMS, SSS, and GSS from colleges and universities in the United States and abroad and cover a wide range of topics related to survey methods, social, and government statistics.

Applied

Yes

Main Sponsor

Survey Research Methods Section

Co Sponsors

Government Statistics Section
Social Statistics Section

Presentations

Bayesian MI-LASSO for Variable Selection on Multiply-Imputed Data: Application to an Environmental Health Survey

Multiple imputation is widely used for handling missing data in surveys. For variable selection on multiply-imputed datasets, however, if selection is performed on each imputed dataset separately, it can result in different sets of selected variables across datasets. MI-LASSO, one of the most commonly used approaches to this problem, regards the same variable across all separate imputed datasets as a group variable and exploits the group LASSO to yield a consistent variable selection across all the multiply-imputed datasets. In this paper, we extend MI-LASSO to a Bayesian framework and propose four Bayesian MI-LASSO models for variable selection on multiply-imputed data, including three shrinkage prior-based and one Spike-and-Slab prior-based methods. To overcome the limitations of traditional, threshold-based model selection, we further develop a principled four-step projection predictive procedure that provides robust variable selection and facilitates coherent post-selection inference. Simulation studies showed that the Bayesian MI-LASSO models outperformed MI-LASSO and other alternative approaches, achieving higher specificity and lower mean squared error across a range of settings. We illustrated these methods through an environmental health survey application using a multiply-imputed dataset from the University of Michigan Dioxin Exposure Study. The R package BMIselect is available on CRAN. 

Keywords

Bayesian models

Environmental health survey

Group LASSO

Multiple imputation

Projection predictive variable selection 

Speaker

Jungang Zou

Co-Author(s)

Sijian Wang
Qixuan Chen, Columbia University

Causal decomposition of Black-white disparities in NIH's selection of proposals for panel discussion

Prior research on NIH R01 grants peer review found that selection of proposals for panel discussion constituted the decision point with the largest contribution to Black-white NIH funding disparities. Although selection into discussion is driven by Preliminary Overall Impact Scores, it is not fully deterministic. In this paper, we use causal analyses on preliminary peer review scores for R01 submissions to evaluate factors contributing to Black-white disparities in selection into discussion. We first use causal decomposition methodology to explore whether hypothetical interventions eliminating Black-white disparities and/or differences in attributes would reduce Black-white disparities in selection of proposals for panel discussion. The attributes we consider include those associated with applicants (degree, career stage, their institution's funding bin) and their submitted proposals (application type, amended status, Preliminary Overall Impact Score, and NIH assigned Administering Organization and Integrated Review Group). Under reasonable assumptions, we find that among the attributes considered, Preliminary Overall Impact Score was the only one that, after equalizing, could dissolve Black-white disparities in selection into discussion. In addition, we use causal decomposition to conduct a thought experiment to study the potential impacts of the NIH's move to binary scoring for the Investigator and Environment criteria in the new Simplified Peer Review Framework on Black-white disparities in selection into discussion. Analyses from our thought experiment suggest that changing to binary scoring of the Investigator-and-Environment factor may not have any effect on disparities in selection into discussion. Overall, under the assumptions of the causal decomposition framework, these results suggest that efforts aimed at reducing disparities in selection into discussion should focus on identifying upstream, policy-relevant, and practically feasible interventions for eliminating racial disparities in Preliminary Overall Impact Scores. In contrast, the transition to the Simplified Peer Review Framework is unlikely to substantially reduce disparities in selection into discussion. 

Keywords

Causal inference

Causal decomposition

Research grant review

Disparity analysis

Policy evaluation

Partial interference 

Speaker

Shreya Prakash

Co-Author(s)

Elena Erosheva, University of Washington
Carole Lee, University of Washington

Characterizing Retrospective Constraints of Survey Designs for Estimating Population-level Causal Effects

Evaluating causal effects in a general population is a critical step for decision makers. Large general use population surveys provide a unique data source for extracting causal evidence. A key limitation is that most surveys are designed primarily for accurate population description rather than causal effect estimation. This paper seeks to characterize some retrospective constraints of survey data for post-hoc causal effect estimation. First, assuming multi-phase selection from a finite population of potential outcomes, we give examples of structural scenarios where construction of the survey weights may be informative for exposure status, leading to biased effect estimates. We then show that a post-hoc weight class adjustment may be applied to reduce bias in this setting. Further, sensitivity analysis relaxing the ignorability assumption is important for observational studies. However, sensitivity models for weighting estimators based on the percentile bootstrap may be anticonservative if the percentile bootstrap is not adapted to the complex survey design. We show that survey bootstrap techniques can properly account for this sampling variation, leading to valid confidence intervals. Simulation studies demonstrate the superior empirical performance of our approach over conventional methods. Finally, we illustrate these concepts using the National Youth and Tobacco Survey to estimate the causal effect of e-cigarette use on the future intention to smoke conventional cigarettes among middle and high school students. 

Keywords

Causal inference

complex survey design

sensitivity analysis

generalizability 

Speaker

Sean Tomlin

Co-Author(s)

Rebecca Andridge, The Ohio State University
Bo Lu, The Ohio State University

Inference with weights: Residualization produces short, valid intervals for varying estimands and varying resampling processes

Weighting procedures are used in observational causal inference to adjust for covariate imbalance within the sample. Common practice for inference is to estimate robust standard errors from a weighted regression of outcome on treatment. However, it is well known that weighting can inflate variance estimates, sometimes significantly, leading to standard errors and confidence intervals that are overly conservative. We instead examine and recommend the use of robust standard errors from a weighted regression that additionally includes the balancing covariates and their interactions with treatment. We show that these standard errors are more precise and asymptotically correct for weights that achieve exact balance under multiple common resampling frameworks, including design-based and model-based inference, as well as superpopulation sampling with a finite sample correction. Gains to precision can be quite significant when the balancing weights adjust for prognostic covariates. For procedures that balance only approximately or in expectation, such as inverse propensity weighting or approximate balancing weights, our proposed method improves precision by reducing residuals through augmentation with the parametric model. We demonstrate our approach through simulation and re-analysis of multiple empirical studies.  

Keywords

causal inference

weighting estimators

inference 

Speaker

Arisa Sadeghpour

Co-Author(s)

Erin Hartman, UC Berkeley
Chad Hazlett, UCLA

Where to Experiment? Site Selection Under Distribution Shift via Optimal Transport and Wasserstein DRO

ow should researchers select experimental sites when the deployment population may differ from observed data? I formulate the problem of experimental site selection as an \textit{optimal transport problem}, developing methods to minimize downstream estimation error by choosing sites that minimize the Wasserstein distance between population and sample covariate distributions. I develop new theoretical upper bounds on PATE and CATE estimation errors, and show that these different objectives lead to different site selection strategies. I extend this approach by using Wasserstein Distributionally Robust Optimization to develop a site selection procedure robust to adversarial perturbations of covariate information: a specific model of distribution shift. I also propose a novel data-driven procedure for selecting the uncertainty radius the Wasserstein DRO problem, which allows the user to benchmark robustness levels against observed variation in their data. Simulation evidence, and a reanalysis of a randomized microcredit experiment in Morocco (Crepon et al.), show that these methods outperform random and stratified sampling of sites when covariates have prognostic $R^2 > .5$, and alternative optimization methods i) for moderate-to-large size problem instances ii) when covariates are moderately informative about treatment effects, and iii) under induced distribution shift. 

Keywords

Causal inference

Optimization

Optimal transport

Distribution Shift 

Speaker

Adam Bouyamourn, Princeton University