Data quality, measurement error, and questionnaire design

Joseph Rodhouse Chair
USDA Economic Research Service (ERS)
 
Wednesday, Aug 5: 2:00 PM - 3:50 PM
6445 
Contributed Papers 
Thomas M. Menino Convention & Exhibition Center 
Room: CC-257A 

Main Sponsor

Survey Research Methods Section

Presentations

Survey Cognitive Burden, Question Difficulty, and Victimization Disclosure

Research on the cognitive burden of violence-related topics across survey formats using web panels is limited; this study addresses this gap. Data were drawn from a web panel, and participants were randomly assigned to one of three conditions: an interleafed format beginning with intimate partner–perpetrated physical violence (IP-PV) items followed by any perpetrator-perpetrated stalking and sexual violence (SV) items; a grouped format in which IP-PV items followed stalking and SV items; or a grouped format in which IP-PV items preceded stalking and SV items. After reporting experiences of stalking, SV, and intimate partner–perpetrated violence (IPV; including stalking, contact SV [CSV], and PV), participants rated perceived burden and navigational difficulty. Associations between perceived cognitive burden and victimization were analyzed across conditions, controlling for demographics. Compared with non-victims, victims of CSV, stalking, or IPV had significantly higher odds of reporting cognitive burden (AOR ≥ 2.2 for women; AOR ≥ 1.5 for men) and difficulty navigating the survey (AOR ≥ 2.1 for women; AOR ≥ 2.8 for men). Format interaction effects were largely not observed. 

Keywords

Survey methodology

Questionnaire formats

Cognitive burden

Probability web panels

Violence victimization

Data collection 

Speaker

Jieru Chen, CDC

Co-Author(s)

Carlos Siordia, CDC
Sharon G. Smith, CDC

Exploring Survey Paradata to Understand Respondent’s Interactions for Better Quality Data

In today's business world, finding ways to improve survey collection data and response rates can be a challenge for many survey data collection agencies. One way to achieve this is to utilize paradata; data that is collected about the survey data collection process. Paradata supports statistically informed evaluation, monitoring and managing of survey collection activities to help improve the survey web instrument usability (Kreuter, 2010). This analysis gives insight into how respondents interact with the survey through exploring respondents' break-off trends, navigation of the survey, participation on certain days of the week, utilization of help resources and completion time. In this study, we analyzed paradata from multiple surveys ranging in size and topics to provide a comprehensive understanding of respondent's survey burden. In this presentation, we will share the best methods to utilize large datasets and to enhance the survey respond rate and survey data quality. 

Keywords

Paradata

Data Collection

Survey Response Rate 

Speaker

Sarah Pfeiff, U.S. Census Bureau

Late Entries, Missing Items: Using Timestamp Paradata to Detect Underreporting in a Diary Survey

Accurate measurement of household food acquisitions is critical for understanding food purchasing behavior and sufficiency, yet underreporting remains a persistent challenge in diary-based surveys. This paper examines whether timestamp paradata from the USDA's National Household Food Acquisition and Purchase Survey (FoodAPS) food log diary can be used to predict underreporting of food acquisitions and items. Specifically, temporal patterns in diary completion, such as long lags between food events and recorded entries, are analyzed to assess whether these behaviors are associated with subsequent admissions of forgotten or omitted acquisitions in a post-diary recall survey. Using linked diary, recall, and paradata records from FoodAPS, empirical relationships are estimated between diary completion timing and indicators of underreporting, including self-reported omissions and discrepances in reported acquisition counts. The results highlight the potential value of paradata as a diagnostic tool for identifying underreporting risk in near real time, with implications for adaptive survey design, targeted prompts, and post-survey adjustment strategies that can enhance data quality. 

Keywords

Paradata

Measurement Error

Underreporting

Diary-Based Surveys

Process Data

Data Quality 

Speaker

Joseph Rodhouse, USDA Economic Research Service (ERS)

Honey, I Forgot the Kids: Underreporting of Children in Probability-Based Web Panels

Household-based surveys experience a systemic undercounting of young children. This is consequential for the measurement of children's health. Research on undercounting has not focused on the effects of rostering and reporting errors in probability-based, web panel surveys. Additionally, little is known about how patterns of undercounting influence estimates of key health indicators. The National Center for Health Statistics Rapid Surveys System utilizes two commercial, probability-based web panels. While it was designed to produce estimates of adults, Round 5 was innovated to produce estimates on the health of children. Panel members provide profile data including the number and ages of children in the household. To evaluate potential underreporting of children due to inaccurate profile data, two samples were drawn from distinct groups and screened for eligibility: main (profile data indicating presence of children) and supplemental (profile data indicating no children). This presentation discusses the characteristics of children identified with the supplemental sample and describes the potential impact on health estimates if these children had been excluded from the study. 

Keywords

Children

Health

Undercount

Rostering

Web Panel Surveys

Rapid Surveys System 

Speaker

Jennifer Rammon, NCHS/CDC

Co-Author(s)

Katherine Irimata, National Center for Health Statistics
James Dahlhamer, National Center for Health Statistics
Jessica Jones, National Center for Health Statistics

Sample Heterogeneity across Recruitment Waves in Respondent-Driven Sampling Studies

Respondent-driven sampling (RDS) leverages social networks to recruit participants. Once seeds are recruited by the research team, participants recruit others through chain referrals, and the sample size accumulates over successive waves. This process may demonstrate a key methodological strength of RDS: particularly how RDS can introduce heterogeneity into a sample.
This study examines participant characteristics across recruitment waves in two distinctive RDS studies: (1) an in-person RDS study that targeted persons who inject drugs in Southeast Michigan; and (2) a web RDS survey that sampled adults with Korean heritage in the U.S. For each study, we selected variables spanning a range of thematic domains (e.g., drug use behavior, immigration history) and summarized them through multidimensional principal components analysis. We then visualized the principal components by wave to illustrate whether RDS introduces sample heterogeneity. Further, we compared the levels and patterns of sample heterogeneity between the two RDS studies. This work is supported by the National Secure Data Service Demonstration project, aiming to strengthen efficiencies in data collection efforts. 

Keywords

Respondent-driven sampling

principal components analysis

hard-to-reach population 

Speaker

Joy Wu

Co-Author(s)

Joy Wu
Brady West, Institute for Social Research
Angelina Lu, University of Michigan
Sunghee Lee, University of Michigan

Heterogeneous latent variable modeling of multivariate outcomes

In biomedical studies, multiple outcomes are often measured on the same individuals, providing complementary views of health status. Jointly analyzing these outcomes and understanding their heterogeneity can provide more robust assessment of disease progression and help guide customized treatment decisions. In this work, we propose to model the heterogeneity in multivariate outcomes by formulating a latent individual frailty and linking it to observed covariates via quantile regression while accounting for outcome-specific variations. We develop an efficient estimation procedure based on the conditional score principle. Our modeling and estimation framework can flexibly accommodate continuous, binary, and categorical outcomes. The proposed method demonstrates favorable asymptotic properties and strong finite-sample performance in numerical studies. 

Keywords

Multivariate outcome

Latent variables

Quantile regression

Conditional score.

Heterogeneity 

Speaker

Yi Liu

Co-Author(s)

Limin Peng, Emory University
John Hanfelt, Emory University