Integrating Surveys and Alternative Data Sources throughout the Data Life Cycle

Don Jang Chair
NORC at The University of Chicago
 
John Finamore Discussant
National Center for Science and Engineering Statistics
 
Don Jang Organizer
NORC at The University of Chicago
 
Tuesday, Aug 4: 4:00 PM - 5:50 PM
1552 
Topic-Contributed Paper Session 
Thomas M. Menino Convention & Exhibition Center 
Room: CC-253A 

Applied

Yes

Main Sponsor

Survey Research Methods Section

Co Sponsors

Government Statistics Section
Social Statistics Section

Presentations

Blended Data for the Advance Monthly Retail Trade Survey (MARTS) State Estimates

Since January 2019, the US Census Bureau has produced an experimental data product of estimated Monthly State Retail Sales by combining Monthly Retail Trade Survey (MRTS) data, administrative data on company payrolls, and third-party retail sales data. Unique data features include (1) multi-state companies report aggregate sales across states to MRTS, not state-specific sales; and (2) some of the third-party data is aggregated across companies within states. In collaboration with the Census Bureau, NORC at the University of Chicago has developed alternative monthly state retail sales estimates. Among the many possible alternative approaches for blending data to create monthly state-level estimates, we focus on three general estimation methodologies: direct estimation, based almost entirely on data from the state and month of interest; model-assisted estimation, which brings in synthetic predictions from a regression model fitted to historical data; and area-level small area estimation (SAE), which models the direct estimates and uses available auxiliary information to borrow strength across industries, states, and months. Preliminary results show that the direct estimates effectively combine the various data sources and that the model-assisted and SAE methods have considerable promise in reducing mean squared error for monthly estimates of state-level retail sales by industry.  

Keywords

Establishment Surveys

Third Party Data 

Speaker

Benjamin Reist, NORC at The University of Chicago

Co-Author(s)

F. Jay Breidt, NORC at The University of Chicago
Taylor Wing, NORC at The University of Chicago
Stephen Kaputa, US Census Bureau
Edwards Kerstin, NORC at the University of Chicago

Improving the Estimates from Web Panel Surveys Using Probability Reference Surveys and Multiple Imputation

To improve the timeliness of data products, survey researchers and practitioners have increasingly
used web surveys and alike to collect information for population health research and dissemination. From the statistical inferential perspective, some of these web panel surveys may have a lack of a well-defined probability sampling structure while other web panel surveys may be probability-based, but may still be subject to high nonresponse and/or coverage errors. Certain statistical adjustments are therefore needed to make proper inferences using such web panel surveys. With a high-quality reference probability survey available, one popular adjustment approach is to create pseudoweights that properly "weight" the web panel survey samples back to the target population underlying the reference survey in order to produce population-weighted estimates of the target of interest. When the variable of interest is collected in the web survey but not in the reference survey, the analytical question can also be framed as a missing data problem. Thus we propose to multiply impute the missing variable on the reference survey combining the information from both data sources. Through various comparisons, we demonstrate that results from different imputation and pseudoweighting models can be compared and used to understand features of the data to aid the analysis. We illustrate the main features and performance of the multiple imputation strategy using a simulation study. We also present a real data analysis based on the web-based Research and Development Survey and the National Health Interview Survey.
 

Keywords

Nonprobability survey

Nonresponse error

Missingness mechanism

Multiple imputation

Propensity score

Pseudoweights 

Speaker

Yulei He, AbbVie

Integrating Federal Surveys: Data Modernization Efforts on NCSES's Workforce Surveys

Like many federal agencies, the National Center for Science and Engineering Statistics (NCSES) within the U.S. National Science Foundation is facing growing challenges in administering its workforce surveys, including declining response rates, rising costs, and reduced operational resources. In response, NCSES and agencies across the federal government are reexamining their survey portfolios to identify opportunities for improving efficiency and effectiveness through innovation. This could be accomplished by shifting from traditional, survey-only data collection methods toward hybrid approaches that incorporate alternative data sources that meet data quality standards while aiming to reduce respondent burden.

Federal initiatives are reinforcing this shift. The Evidence-Based Policymaking Act of 2018 (Public Law 115-435) promotes data sharing for evidence building, while ongoing modernization efforts across agencies emphasize improving survey efficiency, reducing duplicity, and facilitating greater use of alternative data.

This study explores the integration of NCSES's three flagship workforce surveys—the National Training, Education, and Workforce Survey (NTEWS), the National Survey of College Graduates (NSCG), and the Survey of Doctorate Recipients (SDR). The goal is to develop a single, integrated workforce survey that delivers high-quality policy relevant data in a more cost-effective and resource-efficient manner.

In parallel, NCSES is assessing the potential of using alternative data sources to supplement or partially replace data currently collected through these surveys. This presentation will share findings and lessons learned from this initiative, offering insights that may serve as a model for other federal agencies considering similar approaches to survey modernization and data integration. 

Keywords

Data collection

Data linkage

Questionnaire design 

Speaker

May Aydin, NCSES

Co-Author(s)

Gigi Jones, NCSES
Lisa Mirel, National Science Foundation, National Center for Science and Engineering Statistics