Monday, Aug 3: 8:30 AM - 10:20 AM
1370
Invited Paper Session
Thomas M. Menino Convention & Exhibition Center
Room: CC-156A
Applied
Yes
Main Sponsor
Government Statistics Section
Co Sponsors
Committee on Privacy and Confidentiality
Presentations
The National Center for Health Statistics (NCHS) has a long-established data linkage program that links NCHS health survey data with vital and other administrative records to expand the analytic potential of both the survey and administrative data. Integrating survey and administrative data produces rich analytic resources that can be used to support public health surveillance, evidence-based policymaking, and patient-centered outcomes research. However, combining data from multiple sources can also increase disclosure risk. To mitigate this risk, disclosure avoidance strategies can be applied to the input files prior to linkage or the output files post-linkage. The NCHS Data Linkage Program has been exploring the use of privacy enhancing methodologies, such as privacy preserving record linkage (PPRL) to address input privacy and synthetic data generation to address output privacy. This talk will provide an overview of NCHS' implementation and evaluation of selected PPRL methodologies and a describe a pilot project to generate statistically valid synthetic linked data files. The talk will conclude with a discussion of next steps in NCHS' exploration of privacy enhancing techniques.
Keywords
data linkage
privacy enhancing technologies
synthetic data
privacy preserving record linkage
Speaker
Cordell Golden, National Center for Health Statistics (NCHS/CDC)
Applied synthetic data research is having a translational moment. In the last few years, federal statistical agencies have produced numerous pilot synthetic data products such as those discussed in this session, both at individual agencies and through coordinated efforts like those of the National Secure Data Service. As synthetic data matures, and statistical agencies seek to transition these pilots into production, they will encounter a number of methodological and operational challenges. Barriers to adoption frequently stem from difficulties mapping technical methods onto specific synthetic data generation and evaluation problems, working within computational and personnel constraints, and even addressing whether synthetic data is the most appropriate privacy-enhancing technology for a given scenario. Drawing from my work as both as synthetic data open-source software developer and as a contributor to many of the projects within this session, I'll discuss some cross-agency lessons learned about how agency staff can most successfully navigate the frequently hidden decision-making burden that must be overcome to make high-value synthetic data products a reality.
Synthetic data generation can be used to create datasets that do not contain the exact records of the original dataset but instead retain the statistical properties of the original data. The anonymity of the original data is not compromised since synthetic data do not directly correspond with the original values. Creating synthetic data allows for a tiered access model to maximize access to data while ensuring privacy protection. However, producing synthetic data can prove challenging and resource intensive. The federal statistical community has been exploring and, at times, releasing synthetic or partially synthetic datasets for many years. With emerging needs for more access to data, particularly to help train AI/machine learning models, generating synthetic data and generating it in a way that balances privacy and utility is now at the forefront of many conversations. This talk will focus on two synthetic data projects supported through the National Secure Data Service Demonstration Project. The first explores generating synthetic data for a census survey, the Survey of Earned Doctorates, and the second explores generating synthetic data for the National Clinical Cohort Collaborative data (a large real-world high dimensional dataset with over 30 billion rows of data) in a secure super compute environment. The two talks will highlight the use of open-source code for synthetic data generation, the approach being used to assess fidelity to the original data, the opportunities and challenges with synthesizing each source, and lessons learned that can be used to inform future synthetic data generation projects. The talk will conclude with a discussion of future directions on the use of synthetic data for AI/machine learning research and other emerging needs.
Keywords
National Secure Data Service
Confidentiality
Data Access
Speaker
Lisa Mirel, National Science Foundation, National Center for Science and Engineering Statistics