Monday, Aug 3: 2:00 PM - 3:50 PM
6257
Contributed Papers
Thomas M. Menino Convention & Exhibition Center
Room: CC-254B
Main Sponsor
Section on Statistical Learning and Data Science
Presentations
Double Machine Learning (DML) is a principal approach for causal effect estimation in various settings, adapting a wide range of machine learning methods to estimate the nuisance parameters. However, the variation among different algorithms to estimate the variance of DML is not well characterized. We conduct a comprehensive simulation study to compare the coverage probability of DML confidence intervals across different ML methods. In this study, we evaluate coverage probability, bias, and interval length under both model-based and bootstrap-based inference. To do this, we consider a range of learners, including linear models, regularized regression, tree-based methods, and neural networks. In our simulation, data-generating processes for outcome and treatment models include linear and/or nonlinear relationships, interaction terms, and high-dimensional covariates. Our results demonstrate substantial variability in coverage performance across analytical fits and empirical bootstraps, highlighting that learner choice plays a critical role in reliable DML inference. We further investigate the coverage probabilities of DML using a dataset on rural-urban difference among US counties.
Keywords
Double Machine Learning
Causal Inference
Coverage Probability
Simulation Study
Variance Estimation
Bootstrapping
Federated Learning (FL) is a leading framework for training ML and AI models collaboratively across numerous user devices or sensitive databases. We study the key trade-offs among estimation accuracy, privacy constraints, and communication cost across several methods for differentially private (DP) federated training of Empirical risk minimization or M estimators using noisy gradient descent. The two standard methods in the literature are FedAvg and FedSGD. The first simply averages client estimates and can suffer from high bias, while the second aggregates privatized gradients or estimates from clients at each round and can incur high communication cost. Aimed at improving accuracy at a reduced communication cost, we propose FedHybrid, which uses FedSGD starting with an improved initialization, provided for example, by the FedAvg estimator. Finally, we propose FedNewton, which averages local Newton iterations to reduce bias in FedAvg, achieving an estimation accuracy comparable to FedSGD with much fewer communication rounds when the number of clients grows sufficiently slowly with respect to the total sample size. We establish finite sample upper bounds on the mean-squared error rates of the DP versions of these estimators as functions of the number of clients, sample size per client, the privacy budget, and the number of iterations. Our results reveal an important trade-off between improved accuracy and privacy leakage as the number of iterations increases. We further derive a minimax lower bound on the MSE of any iterative private federated procedure that provides a benchmark to assess the optimality gap of these methods. We compare the performance of the methods for training a logistic regression and a convolutional neural network on the computer vision datasets MNIST and CIFAR-10.
Keywords
Federated Learning
M-estimators
Differential Privacy
μ-GDP Privacy
MSE bounds
Minimax Lower Bound
Statistical inference is increasingly performed in federated environments where individual-level data cannot be pooled. However, practical deployment is often constrained by communication costs that make iterative exchange of summary statistics logistically prohibitive. One-shot federated inference, which requires only a single round of communication, is therefore highly desirable but remains challenging in the presence of distributional heterogeneity, or high-dimensional parameters. To address these challenges, we propose a one-shot federated inference framework that reconstructs the pooled likelihood function by aggregating function approximations of site-specific likelihoods. We use neural networks to approximate local likelihood functions and impose Sobolev-norm regularization to jointly control approximation error in both likelihood values and gradients. We establish theoretical results showing how functional approximation error propagates to estimation error of the resulting federated estimator. We demonstrate its effectiveness through likelihood-based inference for logistic regression and stratified Cox models, with an application to a multi-site EHR study.
Keywords
One-shot Federated Inference
Function Approximation
Sobolev Training
With sensitive data collected across various sites, restrictions on data sharing can hinder statistical inference. Recent works in collaborative iterative algorithms like Federated Learning (FL) have demonstrated methods to perform statistical inference with various classical models under this setup. These tasks present a unique challenge of accounting for both the inherent statistical uncertainty of the estimators and the numerical convergence of the iterates. To the best of our knowledge, theoretical analyses till date have considered either online learning frameworks or implicit algorithms for learning. However, in several real life applications, models can only be trained on limited data collected a priori. On the other hand, implicit algorithms incur additional computational cost and generalization gap relative to their explicit counterparts such as stochastic gradient descent (SGD) at each local iteration. In this work, we explore the combined uncertainty of FL iterates under offline learning with limited data, based on the SGD optimization routine. We explore their small sample properties as well as asymptotic behavior, extending them to a collaborative inference framework.
Keywords
Federated Learning
Inference
Small sample properties
Asymptotics
Offline learning
Variance
Identifiability is a foundational concept in statistics, formalizing what can and cannot be learned from data. It underlies classical results on minimax risk lower bounds and fundamental limits in testing and estimation. More recently, work on identifiability has motivated the study of partially identified models, leading to robust methods in econometrics, the social sciences, and genomics. However, the classical notion of identifiability is inherently asymptotic, characterizing learnability under the true data-generating model–which is only available in the infinite-sample limit. Yet there is no general framework for understanding what is identifiable in finite samples. We propose a definition of finite-sample identifiability that generalizes classical identifiability and is asymptotically consistent with it. Using this framework, we extend core theoretical results to finite-sample and partially identified settings, including bounds on minimax risk, testing power, and confidence set size. We illustrate the implications through examples involving zero-inflated Poisson model and nonparametric average treatment effect estimation under imperfect compliance.
Keywords
Identifiability
Random measure
Finite-sample
Total variation distance
Minimax risk bound
Confidence set