Strategies in Creating Community Area Estimates in Health Surveys

Martha McRoy Chair
NORC at the University of Chicago
 
Stanislav Kolenikov Discussant
NORC at The University of Chicago
 
Martha McRoy Organizer
NORC at the University of Chicago
 
Wednesday, Aug 5: 8:30 AM - 10:20 AM
1040 
Invited Paper Session 
Thomas M. Menino Convention & Exhibition Center 
Room: CC-253A 
This invited session highlights innovative strategies used by state and local health surveys to produce community-level estimates that inform health policy and drive local decision-making. As more state health departments invest in localized research, the need for robust and actionable data has grown. Topics will cover small area estimation techniques, advanced modeling approaches, sample design considerations, multi-year data pooling, and disclosure avoidance practices.

Aligned with the JSM 2026 theme, Communities in Action: Advancing Society, this session emphasizes how statistical methods empower communities by providing the evidence base needed to address health disparities, allocate resources effectively, and promote equitable health outcomes. By focusing on sub-state and subpopulation estimation, the session showcases how statistics can directly support community-driven initiatives and policy innovation.

Applied

Yes

Main Sponsor

Survey Research Methods Section

Co Sponsors

Government Statistics Section
Health Policy Statistics Section

Presentations

Designing a State Health Survey for Substate Estimation

NORC has been partnering with the Illinois Department of Public Health (IDPH) to conduct the Healthy Illinois Survey, a multi-mode address-based (ABS) study similar to the Behavioral Risk Factor Surveillance System (BRFSS). The Healthy Illinois (HIL) Survey is designed to examine a broad set of health topics, including access to health services, chronic health conditions, diet, substance abuse, and exposure to violence. One notable feature of HIL is that it will generate estimates for every county, suburban Cook County municipality, Chicago Community area, and ZIP code groupings in areas supported by population density. Consequently, the study will provide higher-resolution data than have been previously available to study topics such as health equity and social determinants of health throughout Illinois. To accomplish this a sub-frame design has been employed to allow for pooling of sample across years paired with stratified systematic sample design employing a geographic sort to support small area estimation down to the ZIP code level. 

Keywords

ABS study

multi-mode

public health 

Speaker

Ned English, NORC at The University of Chicago

Co-Author(s)

Benjamin Reist, NORC at The University of Chicago
Taylor Wing, NORC at The University of Chicago
Amie Conley, NORC at the University of Chicago

Designing for Change: The 25-Year Evolution of the Ohio Medicaid Assessment Survey

State and local health surveys often have to do a lot with limited resources. They are expected to track trends over time and provide detailed, local-level data—sometimes down to counties or neighborhoods. That is a tough balancing act for survey designers. But it also makes these surveys great testing grounds for trying out new survey techniques and statistical methods. The Ohio Medicaid Assessment Survey (OMAS) is a general population survey of residents of the state of Ohio which collects data on health care access, utilization, and characteristics of the population's health. In the over 25 years since first being fielded, in 1998 as a landline random digit dialing (RDD) survey, the OMAS design has needed to adapt and evolve with the times by changing to a cellphone-only RDD survey before transitioning to an address-based mail/web survey utilizing a dual frame that includes administrative Medicaid records in 2023.

In this presentation, we will detail the different experiments of statistical and design methods that have been tested on the OMAS to help ensure that it provides comparable estimates over time and provide high quality estimates at the state and county levels. We will also detail how we use both direct and indirect (e.g., small area estimation) methods to ensure reasonable precision levels at the county level.  

Keywords

State and Local Surveys

Sample design

Address-based sampling (ABS)

Experimental design

Dual frame design

Ohio Medicaid Assessment Survey 

Speaker

Marcus Berzofsky, RTI International

Co-Author

Caroline Scruggs, RTI International

From Medicare Diagnosis to Underlying Dementia Prevalence: Bridging Models for Local Estimation

Dementia is one of the most pressing public health challenges in the United States, yet reliable geographically disaggregated prevalence estimates remain limited. Estimates by age, sex, and small geographic area are essential for planning, resource allocation, and targeted surveillance, but existing data sources are insufficient to measure dementia prevalence reliably below broad national or regional levels.
In this study, we estimate dementia prevalence at the state and county levels by age group, 65–79 and 80+, and sex by integrating information from the Health and Retirement Study, Medicare administrative records, population data, and other publicly available auxiliary sources. Health and Retirement Study data linked to Medicare records support estimation of the relationship between survey-based dementia classification and Medicare diagnosis codes, but the sample size limits direct subnational inference. Medicare-derived administrative data support granular estimation but capture diagnosed dementia rather than true dementia prevalence and may be affected by differential diagnosis and coding patterns across populations.
We address these challenges using extensions of continuation-ratio models, introduced to the small area estimation literature by Slud, Franco, and Hall (2024) and Rein, Franco, et al. (2024), fit within a hierarchical Bayesian framework. The proposed bridging models use the ordered structure of dementia classification and diagnosis states observed in the linked Health and Retirement Study data, calibrate to local Medicare-based diagnosed dementia estimates, and borrow strength across areas using auxiliary administrative and population covariates. The resulting estimates distinguish diagnosed dementia from underlying dementia burden and provide calibrated small area prevalence estimates for policy-relevant geographic and demographic subgroups.
 

Keywords

Small Area Estimation (SAE)

data integration

public health

dementia

survey statistics 

Speaker

Carolina Franco, NORC at The University of Chicago

Co-Author(s)

Carolina Franco, NORC at The University of Chicago
David Rein, NORC at the University of Chicago
Kan Gianattasio, NORC at the University of Chicago
John Wittenborn, NORC at the University of Chicago

Incorporating Nonprobability Data to Expand Web-Based Health Survey Estimates

Nonprobability web-based data sources, such as opt-in panels and convenience samples, have become increasingly available and new data can be collected at a low cost. However, nonprobability data may suffer from bias as they lack a probability sampling structure and may not apply the same level of quality controls used for probability-based data collections. While there are concerns about the reliability of estimates obtained from nonprobability data sources, statistical approaches to integrate nonprobability data with probability surveys offer opportunities to expand the scope of estimates, particularly among small subpopulations, as well as the precision of those estimates.

This presentation considers an application from the Research and Development Survey (RANDS), a web-based health survey conducted by the National Center of Health Statistics. Recent rounds of RANDS have included both probability-based and nonprobability components, with identical surveys administered to each set of respondents. Methodological strategies for combining data from the probability and nonprobability samples to improve and expand health estimates for small domains are explored and empirical findings are used to demonstrate how these blended approaches can reduce bias and improve efficiency. 

Speaker

Katherine Irimata, National Center for Health Statistics

The Challenge of Health Estimates for Population Subgroups and Multiple Geographies in New York City

The Community Health Survey (CHS), conducted annually since 2002, is a representative survey of New York City (NYC) adults and serves as the Health Department's primary source of data on health conditions and behaviors. Despite an annual sample of approximately 10,000 interviews, estimates for specific geographies of varying sizes, such as Neighborhood Health Action Centers (NHACs) and ZIP Code-based areas, and for population subgroups often fail to meet variance and sample size thresholds required for statistical reporting. Since increasing annual sample sizes is rarely feasible for local government surveys, pooling multiple years of data is a common strategy.

Pooling relies on assumptions about temporal consistency and requires careful handling of the survey design. This paper examines the challenges of pooling survey data and compares two approaches: naïve pooling, defined as stacking datasets with simple weight scaling, and design-consistent multi-year pooling, in which weights are recalibrated to a defined multi-year population. Using CHS data from 2021–2022 and 2022–2023, we compare point estimates and standard errors across multiple health indicators and domains.

We find 15–16% of point estimates differ by more than ±10% across approaches, and up to 25% of standard errors differ by more than ±10%, even after applying minimum sample size thresholds. Differences are most pronounced for small domains and outcomes with temporal variability. These findings suggest that pooling is not just a neutral data-processing step but a methodological choice that defines the estimand and shapes the validity of statistical inference, highlighting the need for design-consistent calibration and explicit analytic guidance in applied surveillance settings.

 

Speaker

Ahuva Jacobowitz, NYC Health Department

Co-Author(s)

TASHEMA BHOLANATH, New York City Department of Health and Mental Hygiene
Stephen Immerwahr, NYC Department of Health and Mental Hygiene

Understanding the Health of California's Neighborhoods: Creating and Disseminating Modeled Sub-County Estimates Using the California Health Interview Survey

Large-scale, random population surveys using Address-Based Sampling (ABS) are designed to provide a representative overview of the target population, but how can they be leveraged to produce insights for smaller geographic areas that are smaller than those provided for within the overall sample design? The UCLA Center for Health Policy Research conducts the California Health Interview Survey (CHIS), a population-based omnibus public health survey of the diverse population of California, which provides important information on the health, health behaviors and access to health care services of Californians. Conducted since 2001, CHIS data are used extensively in California in policy development, service planning and research, and the CHIS is recognized and valued nationally as a model population-based health survey. The sample design of the CHIS provides for direct estimates at the county level for most California counties, but users of CHIS data are interested in estimates at smaller levels of geography. Using data from the CHIS, we build statistical models that relate individual-level health outcomes to neighborhood-level predictors from the five-year ACS data. The fitted CHIS-based models are then applied to ACS and Claritas data to generate small-area estimates of key CHIS health indicators for Census tracts, ZIP codes, cities, and legislative districts. This approach leverages both individual and contextual information to produce stable estimates for areas with limited or no survey samples. The modeled estimates are calibrated and validated, and then evaluated for statistical stability before being cleared for release. These estimates are then provided to users through the free data visualization tool AskCHIS Neighborhood Edition (NE). This tool enables users to search for top health topics at granular levels of geography and produce tables and thematic maps for easy visualization. It also allows the user to define custom aggregations of geographies, such as clusters of tracts or ZIP codes, in a map-based interface, and output data visualizations for the user-defined geography in real-time. Additionally, pollution and pesticide data from the CalEnviroScreen is included in the tool, in order to allow users to explore health indicators in the context of environmental factors as well as demographic characteristics. Since its inception in 2014, there have been nearly 50,000 queries on AskCHIS NE. This presentation will provide a brief discussion of the modeling methodology used to create the small-area CHIS estimates, and a demonstration of the key features of the AskCHIS Neighborhood Edition tool. 

Keywords

small area estimation

population health surveys

data dissemination tools 

Speaker

Todd Hughes, UCLA Center for Health Policy Research

Co-Author(s)

Todd Hughes, UCLA Center for Health Policy Research
ZHEYU JIANG, UCLA CENTER FOR HEALTH POLICY RESEARCH
YuChing Yang, UCLA Center for Health Policy Research
Jacob Rosalez, UCLA Center for Health Policy Research
Ninez Ponce, PhD, MPP, UCLA Center for Health Policy Research