Synthetic Data: Applications Across the Federal Statistical System

Melissa Cidade Chair
Bureau of Justice Statistics
 
Michael Hawes Discussant
U.S. Census Bureau
 
Michael Hawes Organizer
U.S. Census Bureau
 
Monday, Aug 3: 8:30 AM - 10:20 AM
1370 
Invited Paper Session 
Thomas M. Menino Convention & Exhibition Center 
Room: CC-156A 

Applied

Yes

Main Sponsor

Government Statistics Section

Co Sponsors

Committee on Privacy and Confidentiality

Presentations

NCHS Data Linkage Program: Exploring Privacy Enhancing Technologies for Data Linkage

The National Center for Health Statistics (NCHS) has a long-established data linkage program that links NCHS health survey data with vital and other administrative records to expand the analytic potential of both the survey and administrative data. Integrating survey and administrative data produces rich analytic resources that can be used to support public health surveillance, evidence-based policymaking, and patient-centered outcomes research. However, combining data from multiple sources can also increase disclosure risk. To mitigate this risk, disclosure avoidance strategies can be applied to the input files prior to linkage or the output files post-linkage. The NCHS Data Linkage Program has been exploring the use of privacy enhancing methodologies, such as privacy preserving record linkage (PPRL) to address input privacy and synthetic data generation to address output privacy. This talk will provide an overview of NCHS' implementation and evaluation of selected PPRL methodologies and a describe a pilot project to generate statistically valid synthetic linked data files. The talk will conclude with a discussion of next steps in NCHS' exploration of privacy enhancing techniques. 

Keywords

data linkage

privacy enhancing technologies

synthetic data

privacy preserving record linkage 

Speaker

Cordell Golden, National Center for Health Statistics (NCHS/CDC)

The Path from Pilot to Product: Operationalizing Synthetic Data Decision-Making for Federal Statistics

Applied synthetic data research is having a translational moment. In the last few years, federal statistical agencies have produced numerous pilot synthetic data products such as those discussed in this session, both at individual agencies and through coordinated efforts like those of the National Secure Data Service. As synthetic data matures, and statistical agencies seek to transition these pilots into production, they will encounter a number of methodological and operational challenges. Barriers to adoption frequently stem from difficulties mapping technical methods onto specific synthetic data generation and evaluation problems, working within computational and personnel constraints, and even addressing whether synthetic data is the most appropriate privacy-enhancing technology for a given scenario. Drawing from my work as both as synthetic data open-source software developer and as a contributor to many of the projects within this session, I'll discuss some cross-agency lessons learned about how agency staff can most successfully navigate the frequently hidden decision-making burden that must be overcome to make high-value synthetic data products a reality. 

Speaker

Jeremy Seeman, U.S. Census Bureau

Steps to Tiered Access: A review of two synthetic data generation projects

Synthetic data generation can be used to create datasets that do not contain the exact records of the original dataset but instead retain the statistical properties of the original data. The anonymity of the original data is not compromised since synthetic data do not directly correspond with the original values. Creating synthetic data allows for a tiered access model to maximize access to data while ensuring privacy protection. However, producing synthetic data can prove challenging and resource intensive. The federal statistical community has been exploring and, at times, releasing synthetic or partially synthetic datasets for many years. With emerging needs for more access to data, particularly to help train AI/machine learning models, generating synthetic data and generating it in a way that balances privacy and utility is now at the forefront of many conversations. This talk will focus on two synthetic data projects supported through the National Secure Data Service Demonstration Project. The first explores generating synthetic data for a census survey, the Survey of Earned Doctorates, and the second explores generating synthetic data for the National Clinical Cohort Collaborative data (a large real-world high dimensional dataset with over 30 billion rows of data) in a secure super compute environment. The two talks will highlight the use of open-source code for synthetic data generation, the approach being used to assess fidelity to the original data, the opportunities and challenges with synthesizing each source, and lessons learned that can be used to inform future synthetic data generation projects. The talk will conclude with a discussion of future directions on the use of synthetic data for AI/machine learning research and other emerging needs.  

Keywords

National Secure Data Service

Confidentiality

Data Access 

Speaker

Lisa Mirel, National Science Foundation, National Center for Science and Engineering Statistics

Presentation

Speaker

Victoria Bryant, Internal Revenue Service