Wednesday, Aug 5: 10:30 AM - 12:20 PM
1515
Topic-Contributed Paper Session
Thomas M. Menino Convention & Exhibition Center
Room: CC-258C
The authors of the five winning papers of the SRMS/SSS/GSS Student Paper Competition to present their research at this session. The papers were selected from over 30 submitted to SRMS, SSS, and GSS from colleges and universities in the United States and abroad and cover a wide range of topics related to survey methods, social, and government statistics.
Applied
Yes
Main Sponsor
Survey Research Methods Section
Co Sponsors
Government Statistics Section
Social Statistics Section
Presentations
Multiple imputation is widely used for handling missing data in surveys. For variable selection on multiply-imputed datasets, however, if selection is performed on each imputed dataset separately, it can result in different sets of selected variables across datasets. MI-LASSO, one of the most commonly used approaches to this problem, regards the same variable across all separate imputed datasets as a group variable and exploits the group LASSO to yield a consistent variable selection across all the multiply-imputed datasets. In this paper, we extend MI-LASSO to a Bayesian framework and propose four Bayesian MI-LASSO models for variable selection on multiply-imputed data, including three shrinkage prior-based and one Spike-and-Slab prior-based methods. To overcome the limitations of traditional, threshold-based model selection, we further develop a principled four-step projection predictive procedure that provides robust variable selection and facilitates coherent post-selection inference. Simulation studies showed that the Bayesian MI-LASSO models outperformed MI-LASSO and other alternative approaches, achieving higher specificity and lower mean squared error across a range of settings. We illustrated these methods through an environmental health survey application using a multiply-imputed dataset from the University of Michigan Dioxin Exposure Study. The R package BMIselect is available on CRAN.
Keywords
Bayesian models
Environmental health survey
Group LASSO
Multiple imputation
Projection predictive variable selection
Prior research on NIH R01 grants peer review found that selection of proposals for panel discussion constituted the decision point with the largest contribution to Black-white NIH funding disparities. Although selection into discussion is driven by Preliminary Overall Impact Scores, it is not fully deterministic. In this paper, we use causal analyses on preliminary peer review scores for R01 submissions to evaluate factors contributing to Black-white disparities in selection into discussion. We first use causal decomposition methodology to explore whether hypothetical interventions eliminating Black-white disparities and/or differences in attributes would reduce Black-white disparities in selection of proposals for panel discussion. The attributes we consider include those associated with applicants (degree, career stage, their institution's funding bin) and their submitted proposals (application type, amended status, Preliminary Overall Impact Score, and NIH assigned Administering Organization and Integrated Review Group). Under reasonable assumptions, we find that among the attributes considered, Preliminary Overall Impact Score was the only one that, after equalizing, could dissolve Black-white disparities in selection into discussion. In addition, we use causal decomposition to conduct a thought experiment to study the potential impacts of the NIH's move to binary scoring for the Investigator and Environment criteria in the new Simplified Peer Review Framework on Black-white disparities in selection into discussion. Analyses from our thought experiment suggest that changing to binary scoring of the Investigator-and-Environment factor may not have any effect on disparities in selection into discussion. Overall, under the assumptions of the causal decomposition framework, these results suggest that efforts aimed at reducing disparities in selection into discussion should focus on identifying upstream, policy-relevant, and practically feasible interventions for eliminating racial disparities in Preliminary Overall Impact Scores. In contrast, the transition to the Simplified Peer Review Framework is unlikely to substantially reduce disparities in selection into discussion.
Keywords
Causal inference
Causal decomposition
Research grant review
Disparity analysis
Policy evaluation
Partial interference
Evaluating causal effects in a general population is a critical step for decision makers. Large general use population surveys provide a unique data source for extracting causal evidence. A key limitation is that most surveys are designed primarily for accurate population description rather than causal effect estimation. This paper seeks to characterize some retrospective constraints of survey data for post-hoc causal effect estimation. First, assuming multi-phase selection from a finite population of potential outcomes, we give examples of structural scenarios where construction of the survey weights may be informative for exposure status, leading to biased effect estimates. We then show that a post-hoc weight class adjustment may be applied to reduce bias in this setting. Further, sensitivity analysis relaxing the ignorability assumption is important for observational studies. However, sensitivity models for weighting estimators based on the percentile bootstrap may be anticonservative if the percentile bootstrap is not adapted to the complex survey design. We show that survey bootstrap techniques can properly account for this sampling variation, leading to valid confidence intervals. Simulation studies demonstrate the superior empirical performance of our approach over conventional methods. Finally, we illustrate these concepts using the National Youth and Tobacco Survey to estimate the causal effect of e-cigarette use on the future intention to smoke conventional cigarettes among middle and high school students.
Keywords
Causal inference
complex survey design
sensitivity analysis
generalizability
Weighting procedures are used in observational causal inference to adjust for covariate imbalance within the sample. Common practice for inference is to estimate robust standard errors from a weighted regression of outcome on treatment. However, it is well known that weighting can inflate variance estimates, sometimes significantly, leading to standard errors and confidence intervals that are overly conservative. We instead examine and recommend the use of robust standard errors from a weighted regression that additionally includes the balancing covariates and their interactions with treatment. We show that these standard errors are more precise and asymptotically correct for weights that achieve exact balance under multiple common resampling frameworks, including design-based and model-based inference, as well as superpopulation sampling with a finite sample correction. Gains to precision can be quite significant when the balancing weights adjust for prognostic covariates. For procedures that balance only approximately or in expectation, such as inverse propensity weighting or approximate balancing weights, our proposed method improves precision by reducing residuals through augmentation with the parametric model. We demonstrate our approach through simulation and re-analysis of multiple empirical studies.
Keywords
causal inference
weighting estimators
inference
ow should researchers select experimental sites when the deployment population may differ from observed data? I formulate the problem of experimental site selection as an \textit{optimal transport problem}, developing methods to minimize downstream estimation error by choosing sites that minimize the Wasserstein distance between population and sample covariate distributions. I develop new theoretical upper bounds on PATE and CATE estimation errors, and show that these different objectives lead to different site selection strategies. I extend this approach by using Wasserstein Distributionally Robust Optimization to develop a site selection procedure robust to adversarial perturbations of covariate information: a specific model of distribution shift. I also propose a novel data-driven procedure for selecting the uncertainty radius the Wasserstein DRO problem, which allows the user to benchmark robustness levels against observed variation in their data. Simulation evidence, and a reanalysis of a randomized microcredit experiment in Morocco (Crepon et al.), show that these methods outperform random and stratified sampling of sites when covariates have prognostic $R^2 > .5$, and alternative optimization methods i) for moderate-to-large size problem instances ii) when covariates are moderately informative about treatment effects, and iii) under induced distribution shift.
Keywords
Causal inference
Optimization
Optimal transport
Distribution Shift