Monday, Aug 3: 2:00 PM - 3:50 PM
1283
Invited Paper Session
Thomas M. Menino Convention & Exhibition Center
Room: CC-156A
Applied
Yes
Main Sponsor
Social Statistics Section
Co Sponsors
Survey Research Methods Section
Presentations
As synthetic data generation becomes an increasingly viable approach for privacy-preserving microdata dissemination, evaluating the trade-offs between privacy protection and analytical utility remains a key challenge. This presentation focuses on the disclosure risk and utility metrics developed for synthetic electronic health record (EHR) data, designed to be both statistically faithful and privacy-safe. Treating a publicly available synthetic EHR dataset as original data, we generate synthetic health data and assess disclosure risk and data utility across a suite of key metrics, such as distributional similarity, correlation preservation, and model fidelity, to quantify these trade-offs. We also demonstrate how visualizing the balance between disclosure risk and data utility provides an intuitive framework for communicating privacy–utility trade-offs to both technical and applied audiences. This work highlights practical lessons learned from developing and evaluating synthetic data at scale, contributing to emerging best practices that align robust privacy protections with actionable data utility.
This talk will discuss the methods used applying differential privacy (DP) to protect the outputs from the Secure Query Service, an interface which allows individuals to query aggregate statistics on linked individuals from the Internal Revenue Service (IRS) Statistics of Income Division. This talk will show how relaxations of DP were utilized to ensure rigorous privacy protections that satisfy the IRS regulations while also maintaining high quality statistics for users.
Keywords
Differential Privacy
Confidentiality
Secure Query Service
Internal Revenue Service
Post-Secondary Outcomes
Taxes and Earnings Data
We introduce a new class of range restricted formal data privacy standards that condition on owner beliefs about sensitive data ranges. By incorporating this additional information, we can provide a stronger privacy guarantee (e.g. an amplification). The range restricted formal privacy standards protect only a subset (or ball) of data values and exclude ranges (or balls) believed to be already publicly known. The privacy standards are designed for the risk-weighted pseudo posterior (model) mechanism (PPM) used to generate synthetic data under an asymptotic Differential (aDP) privacy guarantee. The PPM downweights the likelihood contribution for each record proportionally to its disclosure risk. The PPM is adapted under inclusion of beliefs by adjusting the risk-weighted pseudo likelihood. We introduce two alternative adjustments. The first expresses data owner knowledge of the sensitive range as a probability, $\lambda$, that a datum value drawn from the underlying generating distribution lies \emph{outside} the ball or subspace of values that are sensitive. The portion of each datum likelihood contribution deemed sensitive is then $(1-\lambda) \leq 1$ and is the only portion of the likelihood subject to risk downweighting. The second adjustment encodes knowledge as the difference in probability masses $P(R) \leq 1$ between the edges of the sensitive range, $R$. We use the resulting \emph{conditional} (pseudo) likelihood for a sensitive record, which boosts its worst case tail values away from 0. We compare privacy and utility properties for the PPM under the aDP and range restricted privacy standards.
Keywords
Data Privacy
Synthetic Data
Differential Privacy
Bayesian Hierarchical Modeling
Privacy Amplification
Asymptotic Differential Privacy
Data synthesis as a strategy for broadening access to private and confidential data has seen an
exponential growth in recent years. The hype around synthetic data is mostly triggered by the rise of
modern AI and its hunger for data. In this context, the synthetic data approach also became
increasingly popular in Computer Science with various generative AI models being turned into data
synthesizers. However, the strong focus on algorithms in the Computer Science literature led to what
I call the three black boxes problem. Current discussions at least in the scientific community often do
not think about the data, ignore the purposes for which the data are generated, and typically do not
try to understand how the synthesis model affects downstream analysis tasks. In this talk, I will
discuss some of the challenges arising from this narrow focus on the algorithm.
Keywords
Generative AI
Synthetic
User
Data