Tuesday, Aug 4: 8:30 AM - 10:20 AM
1764
Topic-Contributed Paper Session
Thomas M. Menino Convention & Exhibition Center
Room: CC-108
Applied
Yes
Main Sponsor
Survey Research Methods Section
Co Sponsors
Government Statistics Section
Social Statistics Section
Presentations
Research and development spending, production of skilled workers, and counts of patents and publications are frequent-used indicators of inputs and intermediate activities within each economy or region's science and engineering (S&E) activities. However, for purposes of science and innovation policy some of the most sought-after indicators are those that quantify the outputs and impacts of S&E activity. Software in general and open-source software (OSS) are rapidly growing outputs of S&E activity that contribute to innovation and productivity. Indicators of this output are essential to understand where and how these tools are created.
S&E activity is increasingly embedded in and documented through software tools that are freely available to use, reuse, and modify. These tools are developed by contributors from all sectors of the economy. OSS allows free access to modifiable digital tools that are used for work and leisure and constitutes a part of intangible investment with the qualities of knowledge-based public goods. Despite its widespread use, the extent and impacts of OSS on the economy and innovation remain largely unknown, which may help illustrate aspects of technology diffusion and flow that would enhance science and technology indicators.
Starting in 2022, the National Science Board's Science and Engineering Indicators report began including measures of OSS for government, academic institutions, and businesses, and for countries where contributors reside (NSB 2022). These indicators were created using data collected from GitHub, the largest open source-code hosting platform in the world, and from the federal government's Code.gov, which catalogs OSS projects developed and shared by government agencies.
This presentation updates and extends indicators developed for and presented in the National Science Board's Science and Engineering Indicators report released in 2022. It describes the methodology for developing these indicators and provides a preview of preliminary indicators that are being prepared for the Science and Engineering Indicators report planned for release in 2026. These indicators include the number of open-source software (OSS) repositories created by year, by U.S. economic sector (academia, government, business, non-profit) and by country, as well as network analysis of OSS collaborations across countries.
This paper presents pilot Physical and Monetary Energy Flow Accounts (PEFA and MEFA, respectively) for the United States, covering 2012–2022/2023. These pilot accounts integrate existing economic and energy statistics to track energy flows through the economy in accordance with international accounting standards. The PEFA compiles physical flows using data from EIA, BEA, and BTS, while the MEFA isolates energy-related transactions from BEA's detailed Supply and Use Tables. We demonstrate methods for attributing energy use to industries and adjusting transportation source data from a territory basis to a residency basis.
A key innovation of this paper is the development of state-level monetary energy supply estimates, allocating national production to all 50 states using BEA regional data and EIA electricity statistics. These experimental estimates represent an early step toward subnational natural resource accounts in the U.S.
While balanced and nearly comprehensive at the national scale, the accounts remain preliminary, with areas for future research and refinement including the residency adjustment for truck transportation (PEFA), greater industry detail (national MEFA), and both energy use and intra-state trade (state-level MEFA).
Keywords
Energy
Natural resource accounts
National accounts
Satellite accounts
Regional accounts
A tiered data access approach provides a structured way to offer multiple data access options, allowing users to engage with data at varying levels based on analytical goals and confidentiality risks. In alignment with Section 3582 of the Foundations for Evidence-based Policymaking Act of 2018 and the FY2026 President's Management Agenda priority to eliminate data silos, the federal statistical system has been exploring tiered access options to expand secure access to confidential data while protecting from inappropriate access and use.
One tiered access option, synthetic data generation, creates datasets that do not contain the exact records of the original dataset but do retain the statistical properties of the original data. The anonymity of the original data is not compromised since synthetic data do not directly correspond with the original values. Federal statistical agencies have been releasing synthetic or partially synthetic datasets for many years. With emerging needs for more access to data, particularly to help train AI/machine learning models, generating synthetic data is at the forefront of many tiered access conversations. However, given the potential mosaic effect of multiple data sources, including non-identifying sources, being combined to reveal sensitive information, producing synthetic data in a way that balances privacy and utility can prove challenging and resource intensive.
This talk will focus on how the creation of synthetic data within a tiered access framework has supported the National Secure Data Service (NSDS) demonstration project. The NSDS demonstration project aims to inform a governmentwide effort to strengthen data access infrastructure for data-driven decision making while ensuring protection of confidential data. The talk will provide background on the NSDS demonstration project with a focus on providing tiered data access options through open source synthetic data generation. In addition, the opportunities and challenges with synthesizing data and socializing their use will be highlighted. The talk will conclude with a discussion of future directions on increasing access to data while maintaining privacy protections in a shared service environment.
Keywords
Tiered access
Synthetic data
Privacy-utility tradeoff
Speaker
John Finamore, National Center for Science and Engineering Statistics