Balancing Fairness and Disclosure Limitation when Creating Synthetic Data for Survey Research
Wednesday, Aug 5: 8:50 AM - 9:05 AM
1907
Contributed Papers
Thomas M. Menino Convention & Exhibition Center
The growing demand for open data has intensified the need for release
strategies that protect respondent confidentiality while preserving the
validity of statistical analyses. Partially synthetic data, which
selectively replaces high-risk records with model-generated values, offers
a promising solution. However, when minority group members are
disproportionately flagged as high risk and targeted for synthesis, standard
generative models may distort subgroup-specific statistics, raising fairness
concerns that existing machine learning approaches do not adequately address
within traditional survey statistical frameworks.
We propose two Bayesian log-linear modeling approaches for balancing
fairness and confidentiality in partially synthetic data generation. The
first modifies the prior distribution to penalize deviations between
synthetic and original subgroup outcome rates. The second incorporates
fairness information through fixed offsets in the synthesis model. Both
methods are designed to preserve subgroup-level statistics without
excessively compromising privacy protection. We evaluate the methods using
2022 American Community Survey data from 12 Midwestern states, targeting
health insurance coverage rates across gender, race/ethnicity, and education
subgroups.
Results show that both methods substantially reduce subgroup-level fairness
error relative to the standard baseline, particularly in larger states with
greater sample sizes. Disclosure risk assessment confirms that fairness
improvements can be achieved without materially increasing re-identification
risk. A systematic fairness-privacy trade-off analysis using a weighted
Composite criterion reveals that the relationship between tuning parameters
and optimal performance is state-specific, with the prior-based method
offering a more stable trade-off across diverse settings. Overall, we
provide a generalizable framework for calibrating fairness and privacy in
synthetic data release and offer practical guidance for statistical agencies
seeking to produce datasets that are simultaneously secure, analytically
useful, and fair across diverse population subgroups.
Disclosure limitation
Partially synthetic data
Fairness constraints
Privacy–utility trade-off
American Community Survey
Main Sponsor
Survey Research Methods Section
You have unsaved changes.