Tuesday, Aug 4: 2:00 PM - 3:50 PM
6009
Contributed Papers
Thomas M. Menino Convention & Exhibition Center
Room: CC-209
Main Sponsor
Biometrics Section
Presentations
In Alzheimer's disease research, an important clinical question for individuals still dementia-free at a given follow-up time is how much longer they will remain so. We address this question in the Alzheimer's Disease Neuroimaging Initiative (ADNI), focusing on baseline amyloid status as the exposure. Estimation is challenging because amyloid status is observed rather than randomized, requiring adjustment for confounding, and because time to dementia onset is heterogeneous and heavily right-censored. To address these challenges with clinically interpretable summaries, we focus on quantiles of the residual time to dementia, whose contrasts across amyloid groups quantify how prognosis differs by exposure. We estimate causal contrasts in quantile residual life using a Bayesian nonparametric enriched Dirichlet process mixture model for the joint distribution of event times, exposure, and baseline covariates, with inference via Bayesian g-computation. The approach accommodates ignorable missing baseline covariates through data augmentation, supports inference across clinically relevant landmark times, and allows sensitivity analysis for residual unmeasured confounding. Simulation studies show good performance under complex heterogeneity and heavy censoring. In ADNI, among individuals still dementia-free at relevant landmark times, elevated versus non-elevated baseline amyloid yielded shorter quantiles of remaining dementia-free time, overall and within baseline diagnostic subgroups.
Keywords
Alzheimer's disease
Bayesian nonparametrics
Causal inference
Quantile residual life
Sensitivity analysis
Principal stratification provides a principled framework for causal inference with post-treatment intermediate variables, but identification often relies on strong assumptions such as principal ignorability. Existing methods typically focus on marginal effects and provide limited insight into covariate-dependent heterogeneity.
In this talk, we develop a Bayesian sensitivity analysis framework for conditional principal causal effects, defined within latent principal strata and indexed by pre-treatment covariates. Under treatment ignorability and a structured model for potential outcome means, we show that conditional principal causal effects are identifiable up to a small number of interpretable sensitivity parameters that capture violations of principal ignorability and dependence between potential intermediate outcomes. These parameters enter the estimand through an explicit identification formula, enabling transparent sensitivity analysis without imposing monotonicity or exclusion restrictions.
We combine Bayesian nonparametric models for identifiable components with prior distributions on sensitivity parameters to account for both sampling and sensitivity uncertainty.
Keywords
Principal stratification
Conditional causal effects
Causal effect heterogeneity
Bayesian sensitivity analysis
Bayesian nonparametrics
Speaker
Zihan Zhu, Case Western Reserve Univ - Cleveland, OH
Co-Author
Fan Li, Yale School of Public Health
Causal inference literature has extensively focused on binary treatments, with increasing attention to multi-valued treatments. However, methods for multiple simultaneously assigned treatments are still understudied. This paper introduces two settings: (1) estimating the effects of multiple concurrent treatments of different types (binary, categorical, and continuous) and the effects of treatment interactions, and (2) estimating the average treatment effect across categories of multi-valued regimens. To obtain robust estimates for both settings, we propose a class of methods based on the double machine learning framework. We use machine learning to flexibly model confounding relationships, which can introduce bias in estimating treatment effects due to regularization bias and overfitting. Our methods overcome such bias through Neyman orthogonality and cross-fitting, and are thus well-suited for complex settings with multiple treatments or regimens. We apply the methods to study the effect of three treatments on HIV-related disease. To our knowledge, this work is the first to apply machine learning for robust estimation of interaction effects in the presence of multiple treatments.
Keywords
Causal inference
Machine learning
Multiple treatments
Observational data
Semiparametric model
The causal inference literature has traditionally focused on estimating the mean of the potential outcome, whereas evaluating how a treatment affects the entire outcome distribution can provide additional information in biomedical research. Quantile treatment effects (QTEs) capture such distributional differences, particularly when outcomes are skewed. However, existing approaches for estimating QTEs make distributional assumptions about the outcome and are thus sensitive to model misspecification. Motivated by an HIV study with skewed outcomes, one of which is subject to detection limits, we propose a doubly robust framework for estimating QTEs based on the cumulative probability model (CPM), which is a rank-based, semiparametric linear transformation model. We develop two CPM-based estimation strategies: (1) an inverse-cumulative distribution function(CDF) approach that first estimates the marginal CDF of potential outcomes using the efficient influence function (EIF) and then obtains marginal quantiles via weighted quantile interpolation by inverting the distribution, and (2) a direct approach that solves the EIF of potential marginal quantiles. The proposed estimators are doubly robust and asymptotically normal. We further extend the framework to probability treatment effects (PTEs) and their conditional counterparts. For statistical inference, we investigate several variance estimation procedures, including empirical variance estimators, influence function-based estimators, sandwich estimators, and the nonparametric bootstrap. Simulation studies illustrate that the empirical sandwich estimator and the nonparametric bootstrap provide doubly robust variance estimation with stable finite-sample performance under nuisance-model misspecification. The proposed methods are evaluated through extensive Monte Carlo simulations and illustrated using an HIV data application.
Keywords
Causal inference
Quantile treatment effects
Probability treatment effects
Semiparametric rank-based regression
Variance estimation of doubly robust estimator
Speaker
Hao Wu
Co-Author(s)
Chun Li, Department of Population and Public Health Sciences, University of Southern California
Bryan Shepherd, Vanderbilt University, School of Medicine
Propensity score analysis is a preferred method for addressing confounding bias in nonrandomized observational studies by balancing the distribution of baseline covariates across exposure groups. However, it requires a positivity assumption and does not always yield optimal bias reduction. An alternative approach, prognostic score analysis, removes confounding bias by balancing baseline prognosis rather than exposure assignment. The relative usefulness of prognostic versus propensity score models remains unclear in evidence-based biostatistical practice. We evaluated differences between prognostic and propensity score analyses using real and simulation-based data. We found that prognostic score estimation yields lower bias with higher coverage probability across different scenarios compared to propensity score methods. This study demonstrates that the choice between prognostic and propensity score analyses depends on the study design, the distributions of the outcome and exposure, the type of exposure, and the presence of missing data on exposure and covariates. We also provide a step-by-step guide for selecting and using these methods in observational studies.
Keywords
Propensity score
Prognostic score
Inverse-probability treatment weighting
Survey-weighted analysis
Observational study
Speaker
Alok Dwivedi, University of Missouri School of Medicine, Columbia
Co-Author
Randi Foraker, University of Missouri-Columbia
Large-scale causal mediation analysis performs a large number of mediation analyses to assess potential causal pathways between an exposure and an outcome, e.g., testing each DNA methylation site as a potential mediator in epigenome-wide studies. Traditional mediation analysis approaches typically rely on the strong assumption of no unmeasured confounding between the mediators and the outcome. This assumption is often violated in observational studies where neither the exposure nor the mediators are randomized, leading to potentially biased inference on mediation effects. To address this challenge, we propose a Factor Analysis-based Mediation Analysis (FAMA) framework that corrects for unmeasured confounding in large-scale mediation analysis settings. FAMA integrates factor analysis with methods for handling omitted-variable bias to estimate natural indirect effects in the presence of unmeasured confounding. We establish the theoretical validity of our approach and demonstrate its robust performance in a variety of confounding scenarios through extensive simulation studies. We further apply FAMA to the analysis of the epigenome-wide Normative Aging Study to investigate the medi
Keywords
Causal mediation analysis;
Factor Analysis;
Large-scale inference;
Unmeasured Confounders.
Randomized experiments balance baseline covariates on average, enabling consistent estimation of the average treatment effect (ATE) as well as treatment effects within different subgroups. Yet chance imbalances inevitably arise in practice, and improving covariate balance can enhance efficiency and credibility. We propose a randomization procedure that rerandomizes treatment assignments until a new heterogeneity-targeted balance criterion is met. This criterion incorporates covariate interactions to directly reduce heterogeneity-related imbalance, thereby improving the precision of treatment effect estimation across subgroups. Within a design-based framework, we study inverse probability weighting (IPW) and augmented IPW estimators under rerandomization, establish their asymptotic properties, quantify efficiency gains from heterogeneity-targeted covariate balance, and provide a valid inference. Simulations and a real-data application demonstrate the effectiveness of the proposed design in finite samples. This work broadens the scope of rerandomization beyond ATE, offering a principled framework for more efficient experimental designs to assess treatment effect heterogeneity.
Keywords
Causal inference
Covariate imbalance
Rerandomization
Experimental design
Treatment effect heterogeneity
Randomization-based inference
Speaker
Mufeng Gao, The University of North Carolina at Chapel Hill
Co-Author(s)
Ke Zhu, NCSU and Duke
Michael Hudgens, University of North Carolina at Chapel Hill
Shu Yang, North Carolina State University, Department of Statistics