Small area estimation, Bayesian estimation, and multilevel models

Patrick Joyce Chair
United States Census Bureau
 
Tuesday, Aug 4: 8:30 AM - 10:20 AM
6449 
Contributed Papers 
Thomas M. Menino Convention & Exhibition Center 
Room: CC-107A 

Main Sponsor

Survey Research Methods Section

Presentations

Labour Taxation and In-Work Poverty in Europe: Evidence from the EU Countries

The in-work at-risk-of-poverty rate has become an increasingly relevant issue in Europe, raising concerns about the effectiveness of labour taxation and labour market institutions in ensuring adequate living standards for employed individuals. This paper examines the relationship between labour taxation and in-work poverty in Europe using a macro-level panel dataset covering 25 EU countries over the period 2013–2024. In-work poverty is measured for single individuals without children earning 50%, 67%, and 80% of the average earnings, standard thresholds used to identify low-wage workers across the wage distribution and consistent with Eurostat's framework. To account for unobserved common factors and cross-sectional dependence, the analysis employs the Common Correlated Effects pooled (CCE pooled) estimator (Pesaran, 2006). The empirical framework controls for key labour market and economic characteristics, including part-time employment, temporary and involuntary employment, and labour productivity. The study highlights the importance of considering labour taxation within a macroeconomic and institutional context when designing policies reduce in-work poverty in Europe. 

Keywords

Common Correlated Effects pooled

In-work at-risk-of-poverty rate

Labour market 

Speaker

Matilde Bini, European University of Rome

Co-Author

Lucio Masserini, Department of Economics and Management, University of Pisa

An Integrated Modelling Framework for Estimating Wealth Index for Small Geographical Areas

Accurately assessing economic well-being in developing countries is challenging due to rapid socioeconomic change and the lack of recent census or local-level survey data. Conventional small area estimation methods often rely on outdated census information, while geospatial data, though timely and granular, provide limited predictors and may introduce bias through preprocessing. This motivates the need for modelling strategies that integrate multiple data sources to improve wealth measurement when census data are unavailable or outdated.
We propose an integrated modelling framework that jointly models two contemporaneous sample surveys estimating similar asset wealth indices, while incorporating geospatial covariates as alternative auxiliary information. Challenges arising from differences in data granularity and quality are addressed through a spatial matching procedure. Our results show that the proposed joint model yields more reliable estimates than conventional univariate approaches, reducing errors from reliance on geospatial data. We conclude by discussing implications for small area estimation and extensions to modelling and variance estimation in data-limited settings. 

Keywords

Small area estimation

Asset-based wealth index

Geospatial covariates

Survey data integration

Spatial matching

Joint modelling 

Speaker

Srijeeta Mitra, University of Maryland College Park

Co-Author

Partha Lahiri, University of Maryland-College Park

Model-based estimators using area-level models with clustered effects

Model-based estimators are widely used in small area estimation problems to provide reliable estimates for domains with small sample sizes. When auxiliary information is available only at the area level, area-level models are typically used. We propose a new estimator based on an area-level model with clustered effects. This model allows for heterogeneity in regression coefficients across areas by incorporating clustered coefficients. Pairwise penalties are used to simultaneously identify clusters and estimate parameters. In simulation studies, we compare the performance of the proposed estimator with existing estimators. Additionally, we apply the new estimator to the Forest Inventory and Analysis (FIA) data. 

Keywords

Area-level models

Clustered effects

Penalty functions

Small area estimation 

Speaker

Xin Wang, San Diego State University

Co-Author

Konnor Payne, San Diego State University

Multivariate Weighted Multilevel Modeling of Multi-Domain Survey Scales in a Prospective Study

Multiple outcomes from the same sample are often analyzed separately, ignoring correlations between outcomes. Multivariate multilevel methods (MM) that explicitly model multiple outcomes simultaneously may be more effective, but are rarely used, particularly in complex surveys. We applied MM to evaluate an intervention using an 8-domain survey assessing population health management capabilities of health centers (HCs). Data were nested at three levels: 310 staff across 32 HCs, surveyed at two time points. To address unequal selection probabilities, we used weighted MM (WMM). Four WMMs were specified under varying assumptions about correlation structures across domains and data levels with robust standard errors estimated. Model performance was assessed using information criteria, absolute fit, and prediction measures. All WMMs outperformed combined univariate models. Models accounting for HC- or staff-level correlations demonstrated superior fit and predictive performance, while the most complex model showed evidence of overfitting and high computational burden. The findings highlight the advantages of WMM and demonstrate the first application of a three-level WMM for 8 outcomes. 

Keywords

Multi-domain survey

Weighted multilevel model

Multivariate weighted multilevel model

Multivariate longitudinal data 

Speaker

Wen Wan

Co-Author(s)

Jacob Tanumihardjo, University of Chicago
Mengqi Zhu
Zahra Hosseinian, University of Chicago
Yolanda O'Neal, University of Chicago
Monica Peek, University of Chicago
Marshall Chin, University of Chicago
Donald Hedeker, The University of Chicago

A Bayesian two-fold small area model for estimation of sub-area means when only area-level aggregate

Estimating subarea-level means is difficult when only aggregated area-level totals are observed. We propose a two-fold Bayesian framework that uses hierarchical modeling and covariate driven borrowing of strength to recover latent subarea signals. To incorporate known area-level totals, we introduce a Soft Constraint Theorem that modifies the posterior distribution through a Kullback–Leibler divergence projection. This approach enforces the aggregate constraint while retaining the flexibility of the posterior and avoiding the rigidity of hard constraints. We further study the uncertainty of the Bayesian estimators from a frequentist perspective by estimating bias and variance based on corrected Markov chain Monte Carlo samples. Through simulation studies, we show that the proposed method provides a good approximation to the empirical mean squared error 

Keywords

Small area estimation

Hierarchical Bayesian modeling 

Speaker

Chen Zhao, Division of Biostatistics & Health Data Science, University of Minnesota

Co-Author

J. Sunil Rao

Design-aware Inference for Ordinal Disparity Indices in Complex Surveys

Fields like public health and labor economics utilize ordered categorical outcomes (e.g., self-rated health) to monitor group disparities. However, survey methodology often overlooks inference for ordinal disparity measures under complex sampling. Unlike linear statistics such as means and proportions, ordinal indices based on stochastic dominance are nonlinear and biased, even in simple designs. Multistage, unequal-probability sampling further complicates point and variance estimation, especially when using software intended for linear statistics or assuming independent sampling.
We propose the Design-Aware Ordinal Weighted Likelihood Bootstrap (D–OWLB) for model-based inference on these measures. D–OWLB embeds a primary sampling unit-level rescaled bootstrap within a weighted likelihood estimation framework using Gamma-perturbed replicate weights. This approach accounts for stratification, clustering, and unequal selection probabilities to target the sampling distribution of nonlinear disparity functionals and produce population-level inference based on these functionals. While this article focuses primarily on health applications, the framework can be implemented for a broad range of variables where inference needs to be made from ordinal data.
Simulations show that while design-agnostic bootstraps underestimate variance as complexity increases, D–OWLB aligns closely with design-based benchmarks. Applied to 2023 BRFSS data, D–OWLB provides better-calibrated intervals for self-rated health disparities compared to standard weighted likelihood methods.
 

Keywords

Ordinal health outcomes

Model-based inference

Complex survey design

Bootstrap

Weighted likelihood

Multilevel models 

Speaker

Ujjayini Das, University of Maryland

Reconciliation of Bayes and empirical Bayes interval estimation with application to SAE

Multilevel normal hierarchical models play an important role in developing statistical theory in multiparameter estimation for a wide range of applications. In this article, we propose a new reconciliation framework of the empirical Bayes and hierarchical Bayes approaches for interval estimation of random effects under a two-level normal model. Our framework shows that a second-order efficient empirical Bayes confidence interval, with empirical Bayes coverage error of order O(m^{-3/2}), can also be viewed as a credible interval whose posterior coverage is close to the nominal level, provided a carefully chosen prior-referred to as a 'matching prior'-is placed on the hyperparameters. While existing literature has examined matching priors that reconcile frequentist and Bayesian inference in various settings, this paper is the first to study matching priors with the goal of interval estimation of random effects in a two-level model. We obtain an area-dependent matching prior on the variance component that achieves a proper posterior under mild regularity conditions. The theoretical results in the paper are corroborated through a Monte Carlo simulation study and a real data analysis. 

Keywords

Credible interval

Empirical best linear unbiased prediction

Linear mixed model

Matching prior 

Speaker

Aditi Sen

Co-Author(s)

Masayo Y. Hirose, Institute of Mathematics for Industry, Kyushu University
Partha Lahiri, University of Maryland-College Park