Enhancing Survey Quality Through Response Modeling, Paradata Analytics, and Coverage Frameworks

Jennifer Ortman Chair
US Census Bureau
 
Wednesday, Aug 5: 10:30 AM - 12:20 PM
6046 
Contributed Papers 
Thomas M. Menino Convention & Exhibition Center 
Room: CC-256 

Main Sponsor

Government Statistics Section

Co Sponsors

Survey Research Methods Section

Presentations

Optimizing Census Bureau Survey Contact Strategies

The Census Bureau collects survey responses from millions of households, requiring efficient and effective contact strategies. Balancing operational cost, response rates, and respondent burden is a core challenge. To address this, we leverage contact history data to model the probability of a completed interview. Predictors include number of contacts, contact method, time-of-day, and day-of-week, with features engineered as binary indicators, aggregate counts, and temporal variables.

We explore a range of statistical and machine learning approaches, including regression models, tree-based methods, and unsupervised techniques such as clustering, to identify key patterns in contact behavior and predict completion likelihood. Models are compared using standard validation frameworks and evaluated with metrics such as ROC curves, classification accuracy, and calibration. By analyzing contact patterns across multiple large surveys, we leverage a high volume of cases and rich operational data to uncover strategies that maximize responses while managing costs. The findings provide actionable insights that may help the Census Bureau improve efficiency and increase response rates. 

Keywords

Survey operations

Survey contact strategies

Machine learning

Statistical models 

Speaker

Elizabeth Mahoney, U.S. Census Bureau

Forecasting Survey Response Rates with Sequential Bayesian Modeling

Accurately projecting survey response rates is a critical part of planning and quality assurance for the Census Bureau's field operations. This paper presents a hierarchical Bayesian model that sequentially predicts the final response rate of ongoing surveys using daily cumulative response rate bounds. The model incorporates group-level effects for seasonal trends, yearly shifts, and survey-length variations (10- versus 11-day survey cycles), leveraging historical patterns and recent trends. Each day's posterior distribution of the final response rate serves as the prior for the subsequent day, allowing the model to dynamically update predictions as new data become available. The model is fitted on a ten-year dataset of monthly surveys, evaluating predictive performance using posterior predictive checks, backtesting, and leave-one-out cross-validation. This method offers a practical tool for real-time monitoring and management of survey operations and can be easily adapted to other applications requiring sequential forecasting in uncertain conditions. 

Keywords

Bayesian Modeling

Forecasting

Survey

Response Rates 

Speaker

Anthony Chiado, US Census Bureau

Improving Low Response Score Estimates by Accounting for ACS Sampling Error

The Census Bureau's Planning Database provides statistics from the Decennial Census and the American Community Survey (ACS) 5-year estimates at the tract and block group levels. These files include the Low Response Score (LRS), which is a predicted value of expected mail self-response used to identify areas that may be hard to survey. Because some of the predictors used to create these measures come from ACS estimates, there are concerns that high sampling variance at lower geographic levels may introduce bias and reduce reliability. To address this issue, we use the Public Use Microdata Sample (PUMS) to produce custom estimates and evaluate the impact of ACS sampling error on these predictors. This paper applies recent Census Bureau research to develop a framework that incorporates LRS while accounting for ACS uncertainty through small-area estimation methods such as the Fay-Herriot model. The goal is to improve accuracy in hard-to-survey areas and provide stakeholders with estimates whose sampling errors are closer to the true variability at small geographic levels. 

Keywords

Low Response Score (LRS)

American Community Survey (ACS)

Planning Database

Sampling error

Small area estimation 

Speaker

Ralph Culver III, U.S. Census Bureau

Co-Author(s)

Maranda Pepe, U.S. Census Bureau
Jeffrey Katen, U.S. Census Bureau

Operationalizing the Hard-to-Count Framework: Integrated Methods for Improving Census Coverage

Accurate population enumeration requires the integration of foundational frames, contact strategies, and response mechanisms that perform reliably across heterogeneous populations. As the steward of the Decennial Census, the U.S. Census Bureau faces increasing operational complexity in attempting to achieve complete coverage, particularly for populations that experience structural barriers to enumeration. These challenges are formalized through the Census Bureau's "Hard-To-Count" (HTC) framework, which characterizes enumeration risk along four dimensions: individuals who are hard to interview, hard to persuade, hard to contact, and hard to locate. Although characteristically distinct, these dimensions can co-occur, producing compounded risks of nonresponse, misclassification, and frame error that cannot be effectively addressed through isolated interventions. By combining advances in artificial intelligence, survey methodology, administrative data integration, and geospatial modeling, Reveal and NORC provide evidence that coordinated, data-driven approaches can meaningfully reduce coverage error, improve operational efficiency, and strengthen the statistical foundations. 

Keywords

Hard-to-Count Populations

Administrative Data

Foundational Frames

Survey Methodology

Artificial Intelligence

Geospatial Data 

Speaker

Taylor Wilson, Reveal Global Consulting

Co-Author(s)

Madeline Kelsch, Reveal Global Consulting
Yezzi Lee, Reveal Global Consulting
Martha Stapleton, NORC at The University of Chicago
Ned English, NORC at The University of Chicago

Exploring Best Practices for Inferential Analysis of Paradata-Derived Statistics

Paradata from internet surveys can yield information about respondent behavior that is useful in evaluating questionnaire performance. Because paradata (e.g., timestamps, browser information, and user actions) are not designed with analysis in mind but rather to facilitate instrument functionality, output files can be complex and ill-structured for analytic purposes. For instance, a respondent's paradata record for the time spent interacting with the questionnaire may be hundreds or thousands of lines long, depending on how many actions occurred each session.

Researchers have adapted to the unwieldy nature of paradata analysis; however, it is often limited to descriptive comparisons without the rigor of statistical inference. The atypical data structure can introduce challenges to statistical best practices, complicating the calculation of variance estimates that serve inferential analysis of paradata-derived statistics. In this paper, the authors apply nonparametric methods to these metrics using data from a nationally representative government survey. These results may better inform researchers in maximizing the utility of this valuable data source. 

Keywords

Paradata

Inference

Nonparametric 

Speaker

Luke Larsen, US Census Bureau

Co-Author

Renee Ellis, US Census Bureau

The Anatomy and Evolution of Survey Error

The survey research literature has documented error in surveys from four broad sources, survey coverage, unit non-response, item non-response, and measurement error. We provide new methods to analyze and decompose survey error in an empirical Total Survey Error (TSE) framework that measures error in terms of bias, further subdividing the four sources into false positives, false negatives, and errors in amounts. We apply our approach to comprehensively measure error in the CPS ASEC, the official source of income and poverty statistics in the U.S., using linked administrative records for ten income sources over two decades. Our results show severe and rising underreporting of income and program receipt, driven primarily by measurement error among survey respondents on the extensive margin (false negatives). As a result, improving reporting among respondents may be the most promising margin for remedy. Non-respondent imputations are particularly noisy, with positive and negative errors that tend to offset in net terms. Moreover, we find strong evidence that absolute error (noise) has risen over time. Our approach clarifies the sources of bias and where improvements could be greatest. 

Keywords

Total Survey Error

Measurement Error

Survey Bias 

Speaker

Bruce Meyer, The University of Chicago

Co-Author(s)

Nikolas Mittag, CERGE-EI
Derek Wu, University of Virginia
Anthony Tatarka, University of Wisconsin
Patrick Langetieg, Internal Revenue Service