A Simulation study on coverage probabilities of Double Machine Learning: analytical vs bootstrap CI
Monday, Aug 3: 2:05 PM - 2:20 PM
3702
Contributed Papers
Thomas M. Menino Convention & Exhibition Center
Double Machine Learning (DML) is a principal approach for causal effect estimation in various settings, adapting a wide range of machine learning methods to estimate the nuisance parameters. However, the variation among different algorithms to estimate the variance of DML is not well characterized. We conduct a comprehensive simulation study to compare the coverage probability of DML confidence intervals across different ML methods. In this study, we evaluate coverage probability, bias, and interval length under both model-based and bootstrap-based inference. To do this, we consider a range of learners, including linear models, regularized regression, tree-based methods, and neural networks. In our simulation, data-generating processes for outcome and treatment models include linear and/or nonlinear relationships, interaction terms, and high-dimensional covariates. Our results demonstrate substantial variability in coverage performance across analytical fits and empirical bootstraps, highlighting that learner choice plays a critical role in reliable DML inference. We further investigate the coverage probabilities of DML using a dataset on rural-urban difference among US counties.
Double Machine Learning
Causal Inference
Coverage Probability
Simulation Study
Variance Estimation
Bootstrapping
Main Sponsor
Section on Statistical Learning and Data Science
You have unsaved changes.