Achieving Fairness in AI with Synthetic Data
Wednesday, Aug 5: 8:35 AM - 9:00 AM
Invited Paper Session
Thomas M. Menino Convention & Exhibition Center
Artificial intelligence and machine learning increasingly inform decisions in hiring, lending, healthcare, and justice. Yet real-world datasets often encode historical bias, and models trained on them can reproduce or amplify inequities. Pre-processing via fair synthetic data is a promising,: if we can generate data that mitigates bias at the source while preserving signal, downstream models can be both fair and useful. This talk introduces FDA (Fair synthetic data via Data Augmentation), a statistically principled framework that makes the fairness–faithfulness trade-off explicit and controllable. FDA jointly models a fair submodel and a faithful submodel, coupled by a single parameter $\alpha \in [0,1]$ that quantifies the fraction of bias removed. We prove clear operating points: $\alpha=0$ yields maximal fairness (with larger deviation from the original distribution), $\alpha=1$ recovers the original data in probability (hence in distribution), and intermediate $\alpha$ values guarantee calibrated compromises with interpretable bounds. Practically, FDA samples directly from simple predictive distributions, avoiding heavy black-box training. We further provide theory connecting FDA's $\alpha$ to fairness of downstream models. Together, these results deliver a transparent, efficient, and deployable path to generating fair synthetic data without sacrificing essential statistical structure.
fairness
synthetic data
You have unsaved changes.