A Simulation study on coverage probabilities of Double Machine Learning: analytical vs bootstrap CI

Haozheng Xu Speaker
 
Qingyan Xiang Co-Author
Vanderbilt University Medical Center
 
Siyuan Ma Co-Author
 
Monday, Aug 3: 2:05 PM - 2:20 PM
3702 
Contributed Papers 
Thomas M. Menino Convention & Exhibition Center 
Double Machine Learning (DML) is a principal approach for causal effect estimation in various settings, adapting a wide range of machine learning methods to estimate the nuisance parameters. However, the variation among different algorithms to estimate the variance of DML is not well characterized. We conduct a comprehensive simulation study to compare the coverage probability of DML confidence intervals across different ML methods. In this study, we evaluate coverage probability, bias, and interval length under both model-based and bootstrap-based inference. To do this, we consider a range of learners, including linear models, regularized regression, tree-based methods, and neural networks. In our simulation, data-generating processes for outcome and treatment models include linear and/or nonlinear relationships, interaction terms, and high-dimensional covariates. Our results demonstrate substantial variability in coverage performance across analytical fits and empirical bootstraps, highlighting that learner choice plays a critical role in reliable DML inference. We further investigate the coverage probabilities of DML using a dataset on rural-urban difference among US counties.

Keywords

Double Machine Learning

Causal Inference

Coverage Probability

Simulation Study

Variance Estimation

Bootstrapping 

Main Sponsor

Section on Statistical Learning and Data Science