Federated Learning and Inference

Jiawei Zhang Chair
 
Monday, Aug 3: 2:00 PM - 3:50 PM
6257 
Contributed Papers 
Thomas M. Menino Convention & Exhibition Center 
Room: CC-254B 

Main Sponsor

Section on Statistical Learning and Data Science

Presentations

A Simulation study on coverage probabilities of Double Machine Learning: analytical vs bootstrap CI

Double Machine Learning (DML) is a principal approach for causal effect estimation in various settings, adapting a wide range of machine learning methods to estimate the nuisance parameters. However, the variation among different algorithms to estimate the variance of DML is not well characterized. We conduct a comprehensive simulation study to compare the coverage probability of DML confidence intervals across different ML methods. In this study, we evaluate coverage probability, bias, and interval length under both model-based and bootstrap-based inference. To do this, we consider a range of learners, including linear models, regularized regression, tree-based methods, and neural networks. In our simulation, data-generating processes for outcome and treatment models include linear and/or nonlinear relationships, interaction terms, and high-dimensional covariates. Our results demonstrate substantial variability in coverage performance across analytical fits and empirical bootstraps, highlighting that learner choice plays a critical role in reliable DML inference. We further investigate the coverage probabilities of DML using a dataset on rural-urban difference among US counties. 

Keywords

Double Machine Learning

Causal Inference

Coverage Probability

Simulation Study

Variance Estimation

Bootstrapping 

Speaker

Haozheng Xu

Co-Author(s)

Qingyan Xiang, Vanderbilt University Medical Center
Siyuan Ma

Differentially Private Federated Learning with estimation error bounds

Federated Learning (FL) is a leading framework for training ML and AI models collaboratively across numerous user devices or sensitive databases. We study the key trade-offs among estimation accuracy, privacy constraints, and communication cost across several methods for differentially private (DP) federated training of Empirical risk minimization or M estimators using noisy gradient descent. The two standard methods in the literature are FedAvg and FedSGD. The first simply averages client estimates and can suffer from high bias, while the second aggregates privatized gradients or estimates from clients at each round and can incur high communication cost. Aimed at improving accuracy at a reduced communication cost, we propose FedHybrid, which uses FedSGD starting with an improved initialization, provided for example, by the FedAvg estimator. Finally, we propose FedNewton, which averages local Newton iterations to reduce bias in FedAvg, achieving an estimation accuracy comparable to FedSGD with much fewer communication rounds when the number of clients grows sufficiently slowly with respect to the total sample size. We establish finite sample upper bounds on the mean-squared error rates of the DP versions of these estimators as functions of the number of clients, sample size per client, the privacy budget, and the number of iterations. Our results reveal an important trade-off between improved accuracy and privacy leakage as the number of iterations increases. We further derive a minimax lower bound on the MSE of any iterative private federated procedure that provides a benchmark to assess the optimality gap of these methods. We compare the performance of the methods for training a logistic regression and a convolutional neural network on the computer vision datasets MNIST and CIFAR-10. 

Keywords

Federated Learning

M-estimators

Differential Privacy

μ-GDP Privacy

MSE bounds

Minimax Lower Bound 

Speaker

Xiangni Peng

Co-Author(s)

Arnab Auddy, The Ohio State University
Subhadeep Paul, The Ohio State University

Federated Likelihood-based Inference via Sobolev Neural Approximation

Statistical inference is increasingly performed in federated environments where individual-level data cannot be pooled. However, practical deployment is often constrained by communication costs that make iterative exchange of summary statistics logistically prohibitive. One-shot federated inference, which requires only a single round of communication, is therefore highly desirable but remains challenging in the presence of distributional heterogeneity, or high-dimensional parameters. To address these challenges, we propose a one-shot federated inference framework that reconstructs the pooled likelihood function by aggregating function approximations of site-specific likelihoods. We use neural networks to approximate local likelihood functions and impose Sobolev-norm regularization to jointly control approximation error in both likelihood values and gradients. We establish theoretical results showing how functional approximation error propagates to estimation error of the resulting federated estimator. We demonstrate its effectiveness through likelihood-based inference for logistic regression and stratified Cox models, with an application to a multi-site EHR study. 

Keywords

One-shot Federated Inference

Function Approximation

Sobolev Training 

Speaker

Yue WU, University of Pennsylvania

Co-Author(s)

Huiyuan Wang, University of Pennsylvania
Runze Li, Penn State University
Jianqing Fan, Princeton University
Yong Chen, University of Pennsylvania, Perelman School of Medicine

Privacy Aware Collaborative Inference for GLMs

With sensitive data collected across various sites, restrictions on data sharing can hinder statistical inference. Recent works in collaborative iterative algorithms like Federated Learning (FL) have demonstrated methods to perform statistical inference with various classical models under this setup. These tasks present a unique challenge of accounting for both the inherent statistical uncertainty of the estimators and the numerical convergence of the iterates. To the best of our knowledge, theoretical analyses till date have considered either online learning frameworks or implicit algorithms for learning. However, in several real life applications, models can only be trained on limited data collected a priori. On the other hand, implicit algorithms incur additional computational cost and generalization gap relative to their explicit counterparts such as stochastic gradient descent (SGD) at each local iteration. In this work, we explore the combined uncertainty of FL iterates under offline learning with limited data, based on the SGD optimization routine. We explore their small sample properties as well as asymptotic behavior, extending them to a collaborative inference framework. 

Keywords

Federated Learning

Inference

Small sample properties

Asymptotics

Offline learning

Variance 

Speaker

Bhaskar Ray

Co-Author(s)

Srijan Sengupta, North Carolina State University
Aritra Mitra, North Carolina State University

Weak and Finite Sample Identifiability

Identifiability is a foundational concept in statistics, formalizing what can and cannot be learned from data. It underlies classical results on minimax risk lower bounds and fundamental limits in testing and estimation. More recently, work on identifiability has motivated the study of partially identified models, leading to robust methods in econometrics, the social sciences, and genomics. However, the classical notion of identifiability is inherently asymptotic, characterizing learnability under the true data-generating model–which is only available in the infinite-sample limit. Yet there is no general framework for understanding what is identifiable in finite samples. We propose a definition of finite-sample identifiability that generalizes classical identifiability and is asymptotically consistent with it. Using this framework, we extend core theoretical results to finite-sample and partially identified settings, including bounds on minimax risk, testing power, and confidence set size. We illustrate the implications through examples involving zero-inflated Poisson model and nonparametric average treatment effect estimation under imperfect compliance. 

Keywords

Identifiability

Random measure

Finite-sample

Total variation distance

Minimax risk bound

Confidence set 

Speaker

Won Gu

Co-Author

Justin Silverman, Penn State University