The Challenge of Health Estimates for Population Subgroups and Multiple Geographies in New York City
Wednesday, Aug 5: 9:35 AM - 9:50 AM
Invited Paper Session
Thomas M. Menino Convention & Exhibition Center
The Community Health Survey (CHS), conducted annually since 2002, is a representative survey of New York City (NYC) adults and serves as the Health Department's primary source of data on health conditions and behaviors. Despite an annual sample of approximately 10,000 interviews, estimates for specific geographies of varying sizes, such as Neighborhood Health Action Centers (NHACs) and ZIP Code-based areas, and for population subgroups often fail to meet variance and sample size thresholds required for statistical reporting. Since increasing annual sample sizes is rarely feasible for local government surveys, pooling multiple years of data is a common strategy.
Pooling relies on assumptions about temporal consistency and requires careful handling of the survey design. This paper examines the challenges of pooling survey data and compares two approaches: naïve pooling, defined as stacking datasets with simple weight scaling, and design-consistent multi-year pooling, in which weights are recalibrated to a defined multi-year population. Using CHS data from 2021–2022 and 2022–2023, we compare point estimates and standard errors across multiple health indicators and domains.
We find 15–16% of point estimates differ by more than ±10% across approaches, and up to 25% of standard errors differ by more than ±10%, even after applying minimum sample size thresholds. Differences are most pronounced for small domains and outcomes with temporal variability. These findings suggest that pooling is not just a neutral data-processing step but a methodological choice that defines the estimand and shapes the validity of statistical inference, highlighting the need for design-consistent calibration and explicit analytic guidance in applied surveillance settings.
You have unsaved changes.