Innovations in analysis of complex survey data

Aditi Sen Chair
 
Thursday, Aug 6: 8:30 AM - 10:20 AM
6446 
Contributed Papers 
Thomas M. Menino Convention & Exhibition Center 
Room: CC-257A 

Main Sponsor

Survey Research Methods Section

Presentations

Extending Hierarchical Generalized Transformation Models to Complex Survey Data

Bradley (2022) introduced hierarchical generalized transformation (HGT) models for joint Bayesian analysis of multiple response types by a two-step composite sampler exploiting Diaconis-Ylvisaker (DY) conjugacy. However, HGT assumes simple random sampling, precluding application to complex surveys like the National Survey of Children's Health (NSCH; Ghandour et al., 2018). This study extends HGT using pseudo-likelihood weighting (Pfeffermann, 1993) raising each observation's likelihood to the power of its survey weight. We prove DY conjugacy is preserved across all three data models- the weighted posteriors retain the same DY form, with weights rescaling sufficient statistics (alphaj + Zij becomes alphaj + wi*Zij; kappaj + cj becomes kappaj + wi*c). Thus, preserving computational efficiency, Bradley's Algorithm 1 requires only minimal modification. This study applies survey-weighted HGT to NSCH data, jointly modeling child screen time (continuous), sleep problems (ordinal), and anxiety (binary), and demonstrate that ignoring survey design yields biased estimates and invalid inference. 

Keywords

complex surveys

hierarchical transformation models

Diaconis-Ylvisaker conjugacy

pseudo-likelihood

Bayesian inference

survey weights 

Speaker

Asma Ul Hosna

Co-Author(s)

Md. Niamul Islam Sium, University of Texas at El Paso
Rene Gutierrez Marquez

Generalized Regression Estimation Under Misspecified Sample Design

Classical design-based survey estimation relies on a properly specified sampling design for valid inference. We consider the properties of regression estimation under a misspecified sampling design, in which the nominal and true inclusion probabilities do not necessarily match. This general misspecified sample design setting encompasses many challenges in the current sample survey environment, especially with national statistical agencies such as the Census Bureau. Under this setting, an asymptotic analysis of the regression estimator, an expression of the bias, and an expression of the variance are presented. Further, a consistent variance estimator is derived and an expression which estimates the bias in-part or in-whole is discussed. This later expression may be used as an indicator of the presence of bias due to misspecification by a practitioner. A simulation study is conducted to support the presented theory. 

Keywords

Design-based

Model-assisted

Probability sampling

Survey asymptotics

Survey sampling 

Speaker

Joseph Engmark, University of Maryland

Co-Author

Jean Opsomer, University of Maryland

Issues in causal inference using matching with longitudinal survey data

Matching methods are well-established as a popular approach to inference, with well-developed tools and methods, and effective use in the public health literature. Several problems encountered in population studies of smoking behavior using longitudinal survey data are outlined, including questions of mediation and of adjusting for prognostic variables. In the setting of large-scale complex surveys, additional questions always include how and when to incorporate survey weights, and how to carry out resampling-based approaches to inference. 

Keywords

Matching

causal inference

survey statistics

mediation 

Speaker

Karen Messer, UCSD Division of Biostatistics and Bioiformatics

Co-Author(s)

Jiayu Chen, UC San Diego
Natalie Quach, UC-San Diego

Ranking Tables and Uncertainty

Abstract
There is broad and deep interest in ranking units in a collection of K(≥ 2) units or populations. For
example, estimated rankings of K populations may communicate quickly high level useful messages
regarding traits of those populations with desirable (or undesirable) ranks. We consider the question,
"Should a table showing sample estimates for K populations that is produced by a national
statistical agency be presented explicitly as a ranking table?" Assuming the answer is "Yes", we
discuss a method to help produce such a ranking table which presents an overall estimated ranking
and construct a 100(1-α)% joint con�fidence region for the overall true ranking of the K populations.
Assuming normality for the K independent estimators, we only need K estimates
and their associated standard errors to produce this joint confi�dence region. We also illustrate a
theoretically based visual showing at once: (1) a joint confi�dence region revealing uncertainty in
the estimated ranking; (2) possible true rankings, beyond the estimated ranking; (3) a marginal
con�fidence set for population k true rank, for k=1,...,K; and (4) a marginal confidence set for each rank r. 

Keywords

Estimated Ranking

DIFF Joint Confidence Region

Statistical Agencies 

Speaker

Tommy Wright, US Census Bureau

Symmetric Coefficients of Variation for Binomial Estimates

Dating back decades, the U.S. Bureau of Labor Statistics (BLS) determines the publishability of Current Population Survey (CPS) labor force statistics by the weighted size of the estimation base. Recently, the effectiveness of this publication rule has attenuated as response rates have declined from around 95 percent to less than 70 percent, leading the BLS to seek more robust data quality standards based on coefficients of variation (CVs), as recommended in U.S. Census Bureau Statistical Quality Standards. However, CVs for binomial data, like labor force participation or unemployment rates, are inherently asymmetric, as the CV of a rate p and its complement 1-p can be quite different despite having equal variances. In this paper, we propose the geometric mean of CV(p) and CV(1-p) as a symmetric coefficient of variation for binomial estimates, resulting in publication standards that only depend on design effects and response counts, and demonstrate its relative stability compared to CV(p). 

Keywords

Current Population Survey

coefficient of variation

symmetric CV

binomial

geometric mean

design effect 

Speaker

Connor Doherty, Bureau of Labor Statistics

Co-Author

Justin McIllece, Bureau of Labor Statistics

Virtually Adjusted Naturally Constrained Entropy (VANCE) Estimator via TRUMP Survey Methodology

We propose what we call a Virtually Adjusted Naturally Constrained Entropy (VANCE) estimator for estimating population totals under complex sampling designs. The methodology is grounded in the principle that entropy is an intrinsic measure of information and should remain naturally constrained, while distortions arising from unequal inclusion probabilities, design effects, and population heterogeneity are accommodated through a virtual adjustment mechanism. The approach avoids direct modification of the entropy structure and achieves stabilization through controlled adjustment, thereby preserving information integrity via what we call a Unified Stabilizing Hyperbolic Adjustment (USHA). In addition, TRUMP Cuts are used as tuning devices to enhance efficiency and accuracy. Expressions for bias and mean squared error of the VANCE estimator are derived. Extensive simulation studies show that the VANCE estimator outperforms the competing estimators considered. This work continues a line of research presented at the Joint Statistical Meetings since 2017, with recent proceedings among the most viewed on Zenodo.. 

Keywords

Empirical Log Likelihood Estimates

Complex Survey Designs

System of non-linear equations

Auxiliary Information

Tuning of Design Weights

Jackknifing and TRUMP Cuts 

Speaker

Sarjinder Singh, Texas A&M University-Kingsville

Co-Author

Stephen Sedory, Texas A & M University - Kingsville