Methodological Advances by 2025 COPSS Emerging Leaders

Ana Maria Ortega-Villa Chair
National Institutes of Health
 
Ana Maria Ortega-Villa Organizer
National Institutes of Health
 
Wednesday, Aug 5: 8:30 AM - 10:20 AM
1142 
Invited Paper Session 
Thomas M. Menino Convention & Exhibition Center 
Room: CC-253C 

Applied

No

Main Sponsor

Committee of Presidents of Statistical Societies

Presentations

Achieving Fairness in AI with Synthetic Data

Artificial intelligence and machine learning increasingly inform decisions in hiring, lending, healthcare, and justice. Yet real-world datasets often encode historical bias, and models trained on them can reproduce or amplify inequities. Pre-processing via fair synthetic data is a promising,: if we can generate data that mitigates bias at the source while preserving signal, downstream models can be both fair and useful. This talk introduces FDA (Fair synthetic data via Data Augmentation), a statistically principled framework that makes the fairness–faithfulness trade-off explicit and controllable. FDA jointly models a fair submodel and a faithful submodel, coupled by a single parameter $\alpha \in [0,1]$ that quantifies the fraction of bias removed. We prove clear operating points: $\alpha=0$ yields maximal fairness (with larger deviation from the original distribution), $\alpha=1$ recovers the original data in probability (hence in distribution), and intermediate $\alpha$ values guarantee calibrated compromises with interpretable bounds. Practically, FDA samples directly from simple predictive distributions, avoiding heavy black-box training. We further provide theory connecting FDA's $\alpha$ to fairness of downstream models. Together, these results deliver a transparent, efficient, and deployable path to generating fair synthetic data without sacrificing essential statistical structure.
 

Keywords

fairness

synthetic data 

Speaker

Bei Jiang, University of Alberta

Bayesian Additive Regression Tree Copula Processes for Scalable Distributional Prediction

We show how to construct the implied copula process of response values from a Bayesian additive regression tree (BART) model with prior on the leaf node variances. This copula process, defined on the covariate space, can be paired with any marginal distribution for the dependent variable to construct a flexible distributional BART model. Bayesian inference is performed via Markov chain Monte Carlo on an augmented posterior, where we show that key sampling steps can be realized as those of Chipman et al. (2010), preserving scalability and computational efficiency even though the copula process is high dimensional. The posterior predictive distribution from the copula process model is derived in closed form as the push-forward of the posterior predictive distribution of the underlying BART model with an optimal transport map. Under suitable conditions, we establish posterior consistency for the regression function and posterior means and prove convergence in distribution of the predictive process and conditional expectation. Simulation studies demonstrate improved accuracy of distributional predictions compared to the original BART model and leading benchmarks. Applications to five real datasets with 506 to 515,345 observations and 8 to 90 covariates further highlight the efficacy and scalability of our proposed BART copula process model. 

Keywords

BART

Copula Process

Distributional Regression

Implicit Copula

Transport Map 

Speaker

Nadja Klein, Karlsruhe Institute of Technology

Bayesian graph-informed disease modeling from neuroimaging to EHR data

Neurodegenerative and complex chronic diseases emerge from interactions that span biological networks and populations. I will introduce two Bayesian frameworks that quantify these interactions, from within-brain propagation to across-disease comorbidity. In the first project, aimed at characterizing Tau protein spread along functional networks in the early course of Alzheimer's disease generated from the A4 study, we jointly model tau propagation, functional connectivity structure, and subgroup heterogeneity using cross-sectional data. By integrating graph-constrained infection dynamics with connectivity patterns, the model infers plausible propagation pathways and subgroup-specific infection sequences. In the second project, using UK Biobank EHRs, we represent disease relationships through a latent hypergraph, where each hyperedge captures the higher-order clusters of diseases that share covariate-dependent risk. By uncovering disease hyperedges and their associated risk factors, we characterize how biological and lifestyle factors jointly influence sets of related conditions. To scale posterior inference while preserving uncertainty, we employ amortized variational inference with neural parameterization. Together, these studies yield probabilistic, graph-informed and interpretable views of disease pathology and organization. 

Speaker

Yize Zhao, Yale University

Symmetric Vaccine Efficacy: Interpretable Estimation and Inference for Vaccine Trials

Traditional measures of vaccine efficacy (VE) are inherently asymmetric, constrained above by 1 but unbounded below. As a result, intervals for VE can extend far below zero, making interpretation challenging and sometimes giving a false impression of evidence that a vaccine is harmful when uncertainty is large. This talk proposes symmetric vaccine efficacy (SVE), a bounded and interpretable alternative to VE that maintains desirable statistical properties while resolving these asymmetries. SVE is defined as a symmetric transformation of observed infection proportions, ensuring estimates remain within (-1, 1) and providing a consistent scale for both beneficial and harmful vaccine effects. We derive a variance and confidence interval for SVE, describe its relationship to traditional VE, and illustrate its application in real trial data. In the example, SVE yields more interpretable uncertainty intervals and clearer graphical summaries compared to standard VE. We then demonstrate open-source tools for computing estimates of SVE and corresponding confidence intervals, available in R through the sve package. 

Keywords

vaccine efficacy

infectious disease

clinical trials

Inference 

Speaker

Lucy D'Agostino McGowan, Wake Forest University