Data Collection, Field Effort, Screening, and Survey Costs

Sunghee Lee Chair
University of Michigan
 
Monday, Aug 3: 2:00 PM - 3:50 PM
6443 
Contributed Papers 
Thomas M. Menino Convention & Exhibition Center 
Room: CC-256 

Main Sponsor

Survey Research Methods Section

Presentations

The Importance of Survey Methods and Quantitative Surveys for Fit-for-Purpose Public Statistics

High-quality public statistics are essential for public policies and the exercise of citizenship. Although survey methods are often overlooked, quantitative surveys play a crucial role in generating relevant public statistics. To inform ongoing debate about the need for quantitative surveys in our evolving information society, two recent Brazilian surveys on specific social and economic issues will be presented. One survey collected data on the effects of a tragic dam collapse on labour, income, and subsistence among affected populations in 45 Brazilian cities. The collapse constitutes Brazil's worst environmental tragedy caused by a technological disaster of the mining industry, in which 43 million m3 of iron ore tailings caused environmental damage, polluting 668km of watercourses. The other is a household survey designed to inform public policy in a single town and to provide local area statistics. The city is a pioneer in Brazil in conducting a large-scale survey, interviewing 13,900 households. The value of household surveys is recognised when the data collected are crucial for decision-making and suitable for their intended purpose. 

Keywords

Survey Methods

Household surveys

Public statistics

Environmental disaster

Policy evaluation 

Speaker

Denise Britz Do Silva, SCIENCE - Sociedade para o Desenvolvimento da Pesquisa Científica

Survey Costs: What Measures Are Used in Survey Methodology and Statistics Research?

Survey costs have long been acknowledged as one of the most important constraints on survey designs. Despite this importance, survey costs and their correlates remain largely hidden. When costs are discussed, different measures are used and inputs into the cost measures are inconsistently defined, leading to difficulties in comparing costs across studies. The goal of this paper is to examine how costs are defined in published survey methodological and statistical articles. To do so, this paper content analyzes over 1500 articles in four journals dedicated solely to survey methodology and statistics research. Costs for each of these papers are coded according to the categories of monetary and nonmonetary costs, the cost metric, the data source for costs, and whether the costs reflect fixed or variable costs. Only 10 percent of these articles report any cost information or include costs in a statistical formula. Of this 10 percent, about 95 percent examine variable costs, but there is little similarity in other aspects of cost measurement, including whether the costs are estimated or observed in data, the units of analysis for costs (e.g., total, costs per sampled unit, costs per complete), or whether costs are reported for a single part of a study design or summed over multiple components. Nine common measures of nonmonetary costs are identified, and almost 30 inputs into monetary costs are used in different combinations across studies. Although standardized cost measures may be developed in the future, for now, researchers should clearly define cost measurements and consider analyses of multiple studies within organizations or across partnering organizations to identify associations between costs, survey design features, and error indicators. 

Keywords

Survey Costs

Data collection

Variable costs 

Speaker

Kristen Olson, University of Nebraska-Lincoln

Incentive Experiments with Mail-to-Web and Text-to-Web Post-Election Surveys

Mail-to-web surveys for address-based samples is a well-established methodology, but mail production timelines can be too rigid to boost response rates efficiently. Text messaging, in contrast, can be scaled up quickly, with the drawback of mismatches between address-based records and phone numbers. VPC and CVI have fielded post-election surveys among voters eligible for our GOTV mail since 2023. To ensure response and reach underrepresented groups, we used both recruitment modes along with monetary incentives.

We embedded a randomized experiment into the 2025 survey to test if response rates vary by incentive amounts under each mode. The survey targets included those under 35, people of color, and unmarried women. Some of these groups may be harder to survey, providing a test case to measure effects of incentives on not only rates of response but also its composition. We created two experimental conditions for each mode with these incentives: $0 vs. $5 (text) and $3 vs. $5 (mail). The analysis will present effects on response rates and cost/complete and examine if changing incentives affect potential respondent traits, including demographics and past political participation. 

Keywords

mail-to-web survey

text-to-web survey

survey response rates

survey incentives

underrepresented groups

survey sample composition 

Speaker

Yi Wu, Voter Participation Center / Center for Voter Information

Co-Author(s)

Austin Reed, Angle Mastagni Mathews Political Strategies LLC
Isaiah Bailey, Voter Participation Center / Center for Voter Information
Jenna Zitomer, Voter Participation Center / Center for Voter Information
Karuna Koppula, Voter Participation Center / Center for Voter Information
Tim Lumpkins, Voter Participation Center / Center for Voter Information
Matthew Haney, Voter Participation Center / Center for Voter Information
John Malloy, Voter Participation Center / Center for Voter Information

Decision Rules for Field Effort Reduction in the Decennial Census

The Decennial Census is the largest peacetime mobilization in the United States, including a field operation that requires hundreds of thousands of enumerators that knock on doors of housing units who have not self-responded to the Census. This nonresponse follow-up (NRFU) operation is very costly. In order to control NRFU-related data collection costs and reduce contact burden on the US population, the Census uses internally-held administrative data (AD) to identify housing units that can be enumerated with in-house data, reducing the need for repeated enumerator follow-ups.

A data quality-driven approach to using these AD to enumerate NRFU-eligible housing units could reduce the overall cost of the NRFU operation, while allowing enumerators to focus on cases with unavailable or poor quality AD. The 2026 Census Test will incorporate a decision rules framework to determine which housing units could accurately and sufficiently be enumerated using administrative records, reducing the need for NRFU contact attempts on these cases.

This talk will include a discussion of the overall decision rules framework, the component models, and select simulation results. 

Keywords

adaptive and responsive design

data quality

survey costs

predictive models

bayesian methods 

Speaker

Stephanie Coffey, US Census Bureau

Co-Author

Jennifer Bernard, U.S. Census Bureau

Predicting Unoccupied Addresses for the 2030 Census

For the 2020 Census, the Census Bureau implemented models using administrative records (AR) to identify addresses with a high likelihood of an unoccupied status. These models were deployed during its nonresponse followup (NRFU) operation to limit contact attempts to unoccupied units. By doing so, field interviewers could allocate more time to occupied households that had not responded.

This research analyzes scenarios in which addresses with a high predicted probability of being unoccupied were determined to be occupied during the 2020 NRFU operation. In particular we examine address- and area-level conditions where these differences occurred. We then show improvements to the modeling methodology by incorporating new predictors. We provide general insight into survey methods with respect to how unoccupied addresses can be identified prior to fieldwork. 

Keywords

administrative data

2030 Census 

Speaker

Andrew Keller, US Census Bureau

Too Costly to Screen for Children? How Variables Appended to Address-based Frames Could Help

Household surveys of children typically require large address samples because identifying eligible children relies on household-level screening, which can substantially increase data collection costs. Prior research has shown that the presence-of-children flag appended to the address-based sampling (ABS) frames lacks sufficient accuracy to meaningfully reduce screening costs. Using data from a large national study that screened for children via a mail push-to-web mode, this paper evaluates the feasibility of combining multiple ABS-appended variables to more efficiently identify households with children. Results show that a stratified sampling design can improve screening efficiency and reduce costs, but that effective sample size-accounting for differential sampling rates across strata-is a more appropriate metric than nominal sample size. In addition, the predictive performance of ABS-appended variables varies across children's age groups. The findings also highlight the need to assess the quality of the ABS-appended variables, as improvements over time can directly affect the efficiency gains achievable under a stratified design. 

Keywords

surveys of children

screening efficiency

data collection cost

stratification

data quality 

Speaker

Sipeng Wang, Westat

Co-Author

Daifeng Han, Westat

Reaching Racial and Ethnic Minorities in an Address-based Sample: How Much Efficiency Can We Gain?

Most national surveys aim to produce accurate overall population estimates while supporting comparisons across key sociodemographic groups, including racial and ethnic minorities. This paper evaluates the efficiency and tradeoffs of oversampling racial and ethnic minority groups using the Bayesian Surname–Geography method in the address-based sampling context, in comparison with two alternative approaches based on vendor-appended demographic indicators and proprietary big-data classifiers. Beyond assessing data quality measures such as precision and recall, the paper presents an evaluation framework that explicitly links oversampling choices to design effects arising from unequal selection probabilities. This framework is applied to the empirical data from two national surveys to assess effective sample size gains for both target and non-target populations. The results provide practical guidance for designing oversampling strategies that improve precision for multiple subgroups while avoiding substantial variance inflation in overall estimates. 

Keywords

Oversampling

racial/ethnic minority population

design effect

Bayesian Surname and Geocoding (BSG) 

Speaker

Amy Lin, Westat

Co-Author(s)

Daifeng Han, Westat
Xiaoshu Zhu, Westat
J. Michael Brick, Westat