Advancing Privacy and Usability: Applications for Data-Informed Decisionmaking in Society

Nianqiao Phyllis Ju Chair
Dartmouth College
 
Harrison Quick Discussant
University of Minnesota
 
Jingchen Hu Organizer
Binghamton University
 
Monday, Aug 3: 2:00 PM - 3:50 PM
1283 
Invited Paper Session 
Thomas M. Menino Convention & Exhibition Center 
Room: CC-156A 

Applied

Yes

Main Sponsor

Social Statistics Section

Co Sponsors

Survey Research Methods Section

Presentations

Balancing Privacy and Utility: Evaluating Disclosure Risk and Utility Metrics in Synthetic Health Data

As synthetic data generation becomes an increasingly viable approach for privacy-preserving microdata dissemination, evaluating the trade-offs between privacy protection and analytical utility remains a key challenge. This presentation focuses on the disclosure risk and utility metrics developed for synthetic electronic health record (EHR) data, designed to be both statistically faithful and privacy-safe. Treating a publicly available synthetic EHR dataset as original data, we generate synthetic health data and assess disclosure risk and data utility across a suite of key metrics, such as distributional similarity, correlation preservation, and model fidelity, to quantify these trade-offs. We also demonstrate how visualizing the balance between disclosure risk and data utility provides an intuitive framework for communicating privacy–utility trade-offs to both technical and applied audiences. This work highlights practical lessons learned from developing and evaluating synthetic data at scale, contributing to emerging best practices that align robust privacy protections with actionable data utility. 

Speaker

Minsun Riddles, Westat

Deploying Relaxed Notions of Differential Privacy for the Secure Query Service

This talk will discuss the methods used applying differential privacy (DP) to protect the outputs from the Secure Query Service, an interface which allows individuals to query aggregate statistics on linked individuals from the Internal Revenue Service (IRS) Statistics of Income Division. This talk will show how relaxations of DP were utilized to ensure rigorous privacy protections that satisfy the IRS regulations while also maintaining high quality statistics for users. 

Keywords

Differential Privacy

Confidentiality

Secure Query Service

Internal Revenue Service

Post-Secondary Outcomes

Taxes and Earnings Data 

Speaker

Joshua Snoke, Georgetown University

Privacy Amplification for Synthetic data using Range Restriction

We introduce a new class of range restricted formal data privacy standards that condition on owner beliefs about sensitive data ranges. By incorporating this additional information, we can provide a stronger privacy guarantee (e.g. an amplification). The range restricted formal privacy standards protect only a subset (or ball) of data values and exclude ranges (or balls) believed to be already publicly known. The privacy standards are designed for the risk-weighted pseudo posterior (model) mechanism (PPM) used to generate synthetic data under an asymptotic Differential (aDP) privacy guarantee. The PPM downweights the likelihood contribution for each record proportionally to its disclosure risk. The PPM is adapted under inclusion of beliefs by adjusting the risk-weighted pseudo likelihood. We introduce two alternative adjustments. The first expresses data owner knowledge of the sensitive range as a probability, $\lambda$, that a datum value drawn from the underlying generating distribution lies \emph{outside} the ball or subspace of values that are sensitive. The portion of each datum likelihood contribution deemed sensitive is then $(1-\lambda) \leq 1$ and is the only portion of the likelihood subject to risk downweighting. The second adjustment encodes knowledge as the difference in probability masses $P(R) \leq 1$ between the edges of the sensitive range, $R$. We use the resulting \emph{conditional} (pseudo) likelihood for a sensitive record, which boosts its worst case tail values away from 0. We compare privacy and utility properties for the PPM under the aDP and range restricted privacy standards.  

Keywords

Data Privacy

Synthetic Data

Differential Privacy

Bayesian Hierarchical Modeling

Privacy Amplification

Asymptotic Differential Privacy 

Speaker

Terrance Savitsky, US Bureau of Labor Statistics

Synthetic Data in the Era of AI

Data synthesis as a strategy for broadening access to private and confidential data has seen an
exponential growth in recent years. The hype around synthetic data is mostly triggered by the rise of
modern AI and its hunger for data. In this context, the synthetic data approach also became
increasingly popular in Computer Science with various generative AI models being turned into data
synthesizers. However, the strong focus on algorithms in the Computer Science literature led to what
I call the three black boxes problem. Current discussions at least in the scientific community often do
not think about the data, ignore the purposes for which the data are generated, and typically do not
try to understand how the synthesis model affects downstream analysis tasks. In this talk, I will
discuss some of the challenges arising from this narrow focus on the algorithm. 

Keywords

Generative AI

Synthetic

User

Data 

Speaker

Joerg Drechsler, Institute for Employment Research