Tiered Data Access Frameworks and Synthetic Data Generation
John Finamore
Speaker
National Center for Science and Engineering Statistics
Tuesday, Aug 4: 9:15 AM - 9:35 AM
Topic-Contributed Paper Session
Thomas M. Menino Convention & Exhibition Center
A tiered data access approach provides a structured way to offer multiple data access options, allowing users to engage with data at varying levels based on analytical goals and confidentiality risks. In alignment with Section 3582 of the Foundations for Evidence-based Policymaking Act of 2018 and the FY2026 President's Management Agenda priority to eliminate data silos, the federal statistical system has been exploring tiered access options to expand secure access to confidential data while protecting from inappropriate access and use.
One tiered access option, synthetic data generation, creates datasets that do not contain the exact records of the original dataset but do retain the statistical properties of the original data. The anonymity of the original data is not compromised since synthetic data do not directly correspond with the original values. Federal statistical agencies have been releasing synthetic or partially synthetic datasets for many years. With emerging needs for more access to data, particularly to help train AI/machine learning models, generating synthetic data is at the forefront of many tiered access conversations. However, given the potential mosaic effect of multiple data sources, including non-identifying sources, being combined to reveal sensitive information, producing synthetic data in a way that balances privacy and utility can prove challenging and resource intensive.
This talk will focus on how the creation of synthetic data within a tiered access framework has supported the National Secure Data Service (NSDS) demonstration project. The NSDS demonstration project aims to inform a governmentwide effort to strengthen data access infrastructure for data-driven decision making while ensuring protection of confidential data. The talk will provide background on the NSDS demonstration project with a focus on providing tiered data access options through open source synthetic data generation. In addition, the opportunities and challenges with synthesizing data and socializing their use will be highlighted. The talk will conclude with a discussion of future directions on increasing access to data while maintaining privacy protections in a shared service environment.
Tiered access
Synthetic data
Privacy-utility tradeoff
You have unsaved changes.