Differentially Private Federated Learning with estimation error bounds

Xiangni Peng Speaker
 
Arnab Auddy Co-Author
The Ohio State University
 
Subhadeep Paul Co-Author
The Ohio State University
 
Monday, Aug 3: 2:20 PM - 2:35 PM
3575 
Contributed Papers 
Thomas M. Menino Convention & Exhibition Center 
Federated Learning (FL) is a leading framework for training ML and AI models collaboratively across numerous user devices or sensitive databases. We study the key trade-offs among estimation accuracy, privacy constraints, and communication cost across several methods for differentially private (DP) federated training of Empirical risk minimization or M estimators using noisy gradient descent. The two standard methods in the literature are FedAvg and FedSGD. The first simply averages client estimates and can suffer from high bias, while the second aggregates privatized gradients or estimates from clients at each round and can incur high communication cost. Aimed at improving accuracy at a reduced communication cost, we propose FedHybrid, which uses FedSGD starting with an improved initialization, provided for example, by the FedAvg estimator. Finally, we propose FedNewton, which averages local Newton iterations to reduce bias in FedAvg, achieving an estimation accuracy comparable to FedSGD with much fewer communication rounds when the number of clients grows sufficiently slowly with respect to the total sample size. We establish finite sample upper bounds on the mean-squared error rates of the DP versions of these estimators as functions of the number of clients, sample size per client, the privacy budget, and the number of iterations. Our results reveal an important trade-off between improved accuracy and privacy leakage as the number of iterations increases. We further derive a minimax lower bound on the MSE of any iterative private federated procedure that provides a benchmark to assess the optimality gap of these methods. We compare the performance of the methods for training a logistic regression and a convolutional neural network on the computer vision datasets MNIST and CIFAR-10.

Keywords

Federated Learning

M-estimators

Differential Privacy

μ-GDP Privacy

MSE bounds

Minimax Lower Bound 

Main Sponsor

Section on Statistical Learning and Data Science