Wednesday, Aug 5: 8:30 AM - 10:20 AM
1142
Invited Paper Session
Thomas M. Menino Convention & Exhibition Center
Room: CC-253C
Applied
No
Main Sponsor
Committee of Presidents of Statistical Societies
Presentations
Artificial intelligence and machine learning increasingly inform decisions in hiring, lending, healthcare, and justice. Yet real-world datasets often encode historical bias, and models trained on them can reproduce or amplify inequities. Pre-processing via fair synthetic data is a promising,: if we can generate data that mitigates bias at the source while preserving signal, downstream models can be both fair and useful. This talk introduces FDA (Fair synthetic data via Data Augmentation), a statistically principled framework that makes the fairness–faithfulness trade-off explicit and controllable. FDA jointly models a fair submodel and a faithful submodel, coupled by a single parameter $\alpha \in [0,1]$ that quantifies the fraction of bias removed. We prove clear operating points: $\alpha=0$ yields maximal fairness (with larger deviation from the original distribution), $\alpha=1$ recovers the original data in probability (hence in distribution), and intermediate $\alpha$ values guarantee calibrated compromises with interpretable bounds. Practically, FDA samples directly from simple predictive distributions, avoiding heavy black-box training. We further provide theory connecting FDA's $\alpha$ to fairness of downstream models. Together, these results deliver a transparent, efficient, and deployable path to generating fair synthetic data without sacrificing essential statistical structure.
Keywords
fairness
synthetic data
We show how to construct the implied copula process of response values from a Bayesian additive regression tree (BART) model with prior on the leaf node variances. This copula process, defined on the covariate space, can be paired with any marginal distribution for the dependent variable to construct a flexible distributional BART model. Bayesian inference is performed via Markov chain Monte Carlo on an augmented posterior, where we show that key sampling steps can be realized as those of Chipman et al. (2010), preserving scalability and computational efficiency even though the copula process is high dimensional. The posterior predictive distribution from the copula process model is derived in closed form as the push-forward of the posterior predictive distribution of the underlying BART model with an optimal transport map. Under suitable conditions, we establish posterior consistency for the regression function and posterior means and prove convergence in distribution of the predictive process and conditional expectation. Simulation studies demonstrate improved accuracy of distributional predictions compared to the original BART model and leading benchmarks. Applications to five real datasets with 506 to 515,345 observations and 8 to 90 covariates further highlight the efficacy and scalability of our proposed BART copula process model.
Keywords
BART
Copula Process
Distributional Regression
Implicit Copula
Transport Map
Neurodegenerative and complex chronic diseases emerge from interactions that span biological networks and populations. I will introduce two Bayesian frameworks that quantify these interactions, from within-brain propagation to across-disease comorbidity. In the first project, aimed at characterizing Tau protein spread along functional networks in the early course of Alzheimer's disease generated from the A4 study, we jointly model tau propagation, functional connectivity structure, and subgroup heterogeneity using cross-sectional data. By integrating graph-constrained infection dynamics with connectivity patterns, the model infers plausible propagation pathways and subgroup-specific infection sequences. In the second project, using UK Biobank EHRs, we represent disease relationships through a latent hypergraph, where each hyperedge captures the higher-order clusters of diseases that share covariate-dependent risk. By uncovering disease hyperedges and their associated risk factors, we characterize how biological and lifestyle factors jointly influence sets of related conditions. To scale posterior inference while preserving uncertainty, we employ amortized variational inference with neural parameterization. Together, these studies yield probabilistic, graph-informed and interpretable views of disease pathology and organization.
Traditional measures of vaccine efficacy (VE) are inherently asymmetric, constrained above by 1 but unbounded below. As a result, intervals for VE can extend far below zero, making interpretation challenging and sometimes giving a false impression of evidence that a vaccine is harmful when uncertainty is large. This talk proposes symmetric vaccine efficacy (SVE), a bounded and interpretable alternative to VE that maintains desirable statistical properties while resolving these asymmetries. SVE is defined as a symmetric transformation of observed infection proportions, ensuring estimates remain within (-1, 1) and providing a consistent scale for both beneficial and harmful vaccine effects. We derive a variance and confidence interval for SVE, describe its relationship to traditional VE, and illustrate its application in real trial data. In the example, SVE yields more interpretable uncertainty intervals and clearer graphical summaries compared to standard VE. We then demonstrate open-source tools for computing estimates of SVE and corresponding confidence intervals, available in R through the sve package.
Keywords
vaccine efficacy
infectious disease
clinical trials
Inference