Section on Statistics in Epidemiology 2 - Instrumental Variables & Causal Inference

Jiahao Ping Chair
 
Monday, Aug 3: 10:30 AM - 12:20 PM
6413 
Contributed Papers 
Thomas M. Menino Convention & Exhibition Center 
Room: CC-207 

Main Sponsor

Section on Statistics in Epidemiology

Presentations

A Modified Instrumental Variable Framework for Sensitivity-Based Estimation of Engagement in mHealth

Engagement with mobile health interventions has attracted growing attention, with identifying thresholds for effective engagement as a primary focus. However, in these settings, the exclusion restriction assumption is often unreasonable. Previous methodological advances have addressed this challenge by modifying instrumental variable (IV) approaches with sensitivity analyses to assess how engagement influences outcomes when the exclusion restriction may fail, but these methods have largely been limited to two treatment arms and continuous engagement measures.

We propose a modified instrumental variable (MIV) approach that extends prior work and broadens its application to real-world interventions. Our method accommodates multiple treatment groups and alternative engagement distributions, thereby expanding the range of settings in which engagement effects can be assessed. Under correct specification of a sensitivity parameter, the class of local average treatment effects (LATEs) is identifiable. Estimation proceeds via two-stage least squares to recover a weighted average of LATEs, and simulations demonstrate accurate recovery across diverse treatment and engagement settings. 

Keywords

Instrumental Variables

Sensitivity Analysis

Local Average Treatment Effecs

Multivalued Treatments

Mobile Health 

Speaker

Alexis Schwartz, Vanderbilt University Medical Center

Co-Author

Andrew Spieker, Vanderbilt University Medical Center

Causal Structure–Informed Covariate Adjustment for Debiased Effect Estimation

Causal effect estimation requires principled covariate adjustment to avoid bias induced by inappropriate conditioning. However, in practice, this task is challenging due to inadvertent collider adjustment and linear confounding violation. To address these, we propose a two-step approach that integrates causal structure learning with debiased estimation. In the first step, we perform structural equation modeling (SEM) to detect a directed acyclic graph (DAG) among outcome, exposure and covariates. This allows us to distinguish confounders from colliders based on the inferred causal structure, thus yielding an adjustment set. In the second step, we employ double machine learning (DML) to model confounding relationships and debias the causal effect estimation with the confounding selection in Step 1. The proposed framework is capable of handling colliders through DAG-based structure learning and accommodating nonlinear relationships via nonlinear SEM in the DAG discovery phase and multilayer perceptron in DML. We illustrate the finite-sample performance of this method through extensive simulation and the real-world data from a community-based cohort Framingham Heart Study. 

Keywords

Causal inference

Observational study

Double machine learning

Causal structure learning

Variable selection 

Speaker

Musong Gao

Co-Author(s)

Wei Jin, Boston University
Ching-Ti Liu, Boston University School of Public Health

Doubly Robust Estimation of Desirability of Outcome Ranking (DOOR) Probability

Covariate adjustment is an important tool in medical research for observational studies, and even clinical trial data, since the addition of regression models could improve precision by incorporating imbalanced covariates, and thus help make correct inference. Desirability of outcome ranking (DOOR) is a patient-centric benefit-risk evaluation methodology designed for randomized clinical trials. Still, robust covariate adjustment methods could further expand the compatibility of this method. In traditional DOOR analysis, each participant's outcome is ranked based on pre-specified clinical criteria, where the most desirable rank represents a good outcome with no side effects and the least desirable rank is the worst possible clinical outcome. We develop a causal framework for estimating the population-level DOOR probability, via inverse probability of treatment weighting method, G-Computation method and a Doubly Robust method that combines both. The performance of the proposed methodologies is examined through simulations. We also perform a causal analysis to Multi-Drug Resistant Organism (MDRO) network within Antibacterial Resistant Leadership Group (ARLG). 

Keywords

Causal Inference

G-computation

Infectious Disease

Multi-Drug Resistance

Inverse Probability Weighting

Observational Study 

Speaker

Shiyu Shu, The George Washington University

Co-Author(s)

Toshimitsu Hamasaki, George Washington University Biostatitics Center
Scott Evans, George Washington University
Guoqing Diao, George Washington University

Efficiency Properties of Bias-Corrected Matching Estimators: The Impact of Differing Covariate Roles

Matching estimators are widely used to estimate the Average Treatment Effect on the Treated (ATT). However, their precision depends on the choice of covariates. In particular the role of prognostic covariates (Xp) is less well studied relative to confounders (Xc). In this study, we consider a setting where variables Xc satisfy the unconfoundedness assumption, while variables Xp provide additional prognostic information. We compare several designs including in which outcome relationships are modeled using : (1)Xc only, and (2)both Xc and Xp. Drawing on Hahn's (1998, 2004) semiparametric efficiency bounds, we study the restricted propensity score and demonstrate the potential efficiency gains of the second design. We further examine how a bias-corrected matching estimator performs in finite samples using these designs. This approach decouples assignment sufficiency from prediction sufficiency, allowing prognostic information to reduce variance without introducing propensity-score noise. Simulations confirm that bias correction is essential and that incorporating Xp leads to significant variance reduction. Finally we aim to illustrate this method through a tobacco initiation analysis. 

Keywords

Causal inference

Average Treatment Effect (ATT)

Matching

Semiparametric Efficiency

Bias correction

Prognostic covariates and confounders 

Speaker

Jiayu Chen, UC San Diego

Co-Author

Karen Messer, UCSD Division of Biostatistics and Bioiformatics

Nonparametric Inference with an Instrument under a Separable Binary Treatment Choice Model

Instrumental variable (IV) methods are widely used to infer causal effects in the presence of unmeasured confounding. In this paper, we propose nonparametric inference with an IV under a separable binary treatment choice model, which posits that the odds of the probability of taking the treatment, conditional on the instrument and the treatment-free potential outcome, factor into separable components for each variable. Our approach employs a new variationally independent parameterization based on nuisance functions defined directly from the observed data. This parameterization, coupled with a novel fixed-point argument, enables the use of modern machine learning methods for estimation. We characterize the semiparametric efficiency bound for any smooth functional of the treatment-free potential outcome among the treated and construct a corresponding semiparametric efficient estimator without imposing any unnecessary restriction on nuisance functions. We also describe a generative model and derive empirically falsifiable implications to help assess our assumptions. Our approach extends to nonlinear effects, population-level effects, and nonignorable missing data settings. 

Keywords

Causal Inference

Instrumental Variable

Missing Data

Odds Ratio

Semiparametric Efficiency 

Speaker

Chan Park, University of Illinois Urbana-Champaign

Co-Author

Eric Tchetgen Tchetgen, University of Pennsylvania

Proximal Causal Inference for Interventional Indirect Effects under Intermediate Confounding

Unmeasured confounding remains a major challenge in causal mediation analysis, particularly when treatment affects intermediate variables that subsequently confound the mediator–outcome relationship. In such settings, traditional mediation methods generally fail unless strong no-unmeasured-confounding assumptions are imposed. We propose a nonparametric proximal identification framework using proxy-based adjustment to assess interventional indirect effects (IIEs) under simultaneous unmeasured and intermediate confounding. By extending the proximal g-formula to nested counterfactuals, we derive identification conditions through a sequence of bridge functions-solutions to Fredholm integral equations linking observed proxies to latent factors. We develop a triply robust, locally efficient estimator for the proximal IIE that generalizes semiparametric approaches for complex multivariate systems. Simulation studies demonstrate that our estimator remains unbiased in scenarios where standard methods ignoring unmeasured confounding exhibit significant bias. 

Keywords

Causal Mediation Analysis

Intermediate Confounding

Interventional Indirect Effects

Proximal Causal Inference

Triply Robust Estimation 

Speaker

An-Shun Tai, Institute of Statistics and Data Science, National Tsing Hua University

Regression-based doubly robust estimation of optimal dynamic treatment regimes for binary outcomes

A dynamic treatment regime (DTR) is a sequence of treatment rules, each formulated as a function of a patient's treatment and covariate history, which recommends the next treatment. An optimal DTR can be estimated by backward induction using a sequence of outcome regression models. Since it is difficult to correctly specify outcome models across stages, double robustness for estimating the optimal DTR is a practically useful property. For binary outcomes, doubly robust estimators for multi-stage optimal DTRs have not been well developed because there is a fundamental trade-off in the choice of the link function for the blip model. Most existing methods employ a logit link function; however, because of its non-collapsibility, pseudo-outcomes cannot be constructed solely from the observed outcomes and the estimated blip parameters, hindering doubly robust estimation in multi-stage settings. Although this problem can be circumvented by using collapsible link functions, the resulting outcome mean under the estimated DTR may violate the natural bounds of a binary outcome. We propose a framework that employs doubly robust g-estimation for collapsible measures and a log-odds product for the nuisance treatment-free outcome model to prevent boundary violations of the estimated outcome mean. Simulation studies demonstrate that the proposed method outperforms the existing sequential regression approach (Q-learning) on some performance measures under complex outcome-generating mechanisms. We further illustrate the practical feasibility and utility of our approach by applying it to a multi-stage problem of e-cigarette use and smoking cessation using longitudinal epidemiological data. 

Keywords

Causal inference

Double robustness

Dynamic treatment regimes

G-estimation

Longitudinal data

Regression model 

Speaker

Koshiro Arai

Co-Author

Tomohiro Shinozaki, The University of Tokyo