Sampling, Weighting, Missing Data and Variance Estimation

Jungang Zou Chair
 
Monday, Aug 3: 8:30 AM - 10:20 AM
6448 
Contributed Papers 
Thomas M. Menino Convention & Exhibition Center 
Room: CC-156B 

Main Sponsor

Survey Research Methods Section

Presentations

The Obstetric Rescue Clock: A Survey-Weighted, Doubly Robust Causal Framework for Measuring Delay to Definitive Rescue in High-Risk Delivery Hospitalizations

Maternal mortality and severe maternal morbidity remain critical indicators of obstetric-system performance. Many maternal deaths occur after a sequence of recognizable deterioration, delayed escalation, delayed mobilization of procedural resources, and incomplete or late definitive treatment. Yet national surveillance systems often quantify adverse outcomes after they occur rather than measuring the temporal pathway between hospital admission and definitive rescue. This creates a statistical and operational gap: delay is repeatedly identified in maternal mortality review, but it is rarely defined as a reproducible estimand that can be evaluated using national inpatient data. We developed the Obstetric Rescue Clock, a procedure-day-based framework for measuring whether earlier initiation of rescue-level intervention is associated with lower maternal mortality and resource burden among high-risk delivery hospitalizations.

We conducted a retrospective, complex survey-weighted cohort study using 2021 HCUP NIS data, emulating a target trial of early versus delayed rescue among high-risk delivery hospitalizations with documented rescue procedures. Delivery hospitalizations were identified from the NIS core file and restricted to a prespecified high-risk obstetric phenotype defined by ICD-10-CM diagnosis-prefix groups representing rescue-relevant complications, including hypertensive disorders of pregnancy, hemorrhage or placental complications, obstetric infection or sepsis-related diagnoses, thromboembolic or embolic conditions, respiratory compromise, renal compromise, and peripartum cardiomyopathy. Qualifying rescue procedures were identified using ICD-10-PCS procedure-prefix groups representing transfusion or blood products, uterine or gynecologic definitive procedures, airway or ventilatory support, respiratory support, and vascular occlusion or embolization.

For each hospitalization, time to rescue was defined as the earliest procedure day among qualifying rescue procedures. Let T_i^R = min_{k in P_i^R} PRDAY_{ik}. Early rescue was first qualifying rescue on hospital day 0-1; delayed rescue was first qualifying rescue on hospital day >=2. We specified a target-trial analogue in which rescued high-risk delivery hospitalizations were assigned to early versus delayed rescue and followed through discharge. The estimand was the survey-population average treatment effect among rescued high-risk delivery hospitalizations. Identification required consistency, conditional exchangeability given measured covariates, positivity, and correct incorporation of the NIS survey design.

The primary outcome was in-hospital maternal mortality. Secondary outcomes were length of stay and total hospital charges. We specified a target-trial analogue in which high-risk delivery hospitalizations receiving rescue procedures were assigned to early versus delayed rescue, with follow-up through discharge. The target estimand was the survey-population average treatment effect of early versus delayed rescue among rescued high-risk delivery hospitalizations. For mortality, this estimand is an absolute risk difference; for length of stay and total charges, it is a mean difference. Identification required consistency, conditional exchangeability given measured covariates, positivity, and correct incorporation of the NIS survey design. The interpretation was restricted to timing among hospitalizations with observed rescue procedures; the analysis does not estimate the effect of receiving rescue versus not receiving rescue.

We estimated effects using cross-fit augmented inverse probability weighting. The propensity score g(W)=P(A=1|W) and outcome regressions Q_a(W)=E(Y|A=a,W) were estimated using regularized nuisance models with demographic, socioeconomic, payer, admission, hospital-design, and diagnosis-derived clinical features. The AIPW estimator combined outcome-regression and inverse-probability components and was weighted by NIS discharge weights, providing double robustness under standard assumptions if either the treatment model or outcome model is correctly specified. Cross-fitting reduced overfitting bias in nuisance estimation. Diagnostics included propensity-score overlap, covariate balance using standardized mean differences, missingness summaries, and timing-field checks.

The 2021 NIS core file contained 6,666,752 unweighted inpatient discharges. Among these, 697,552 delivery hospitalizations met the prespecified high-risk obstetric diagnosis screen. A total of 58,605 high-risk delivery hospitalizations had at least one qualifying rescue procedure. Rescue timing was computable for 56,392 hospitalizations, representing 96.2% of those with rescue procedures. These 56,392 hospitalizations formed the analytic cohort. Early rescue occurred in 49,186 hospitalizations, representing 87.2% of the analytic cohort, while delayed rescue occurred in 7,206 hospitalizations, representing 12.8%. The median earliest rescue day was 0, with interquartile range 0–1, indicating that most rescue procedures occurred on the admission day or following day, with a smaller but clinically important delayed tail.

Baseline age was similar across timing groups. The survey-weighted mean age was 30.03 years among early-rescue hospitalizations and 29.56 years among delayed-rescue hospitalizations, with 95% confidence intervals of 29.87–30.20 and 29.36–29.76 years, respectively. Differences in admission, payer, race/ethnicity, emergency department, and elective-admission covariates supported the need for adjustment rather than relying on crude contrasts.

In survey-weighted descriptive analysis, in-hospital mortality was 0.108% among early-rescue hospitalizations, 95% CI 0.078%–0.137%, compared with 0.513% among delayed-rescue hospitalizations, 95% CI 0.345%–0.682%. The descriptive absolute difference was approximately 0.405 percentage points, or 4.05 additional deaths per 1,000 rescued high-risk delivery hospitalizations in the delayed-rescue group. Mean length of stay was 3.105 days after early rescue, 95% CI 3.074–3.136, compared with 6.650 days after delayed rescue, 95% CI 6.443–6.857. Mean total charges were $34,304 after early rescue, 95% CI $32,995–$35,613, compared with $70,665 after delayed rescue, 95% CI $66,135–$75,194. Thus, before doubly robust adjustment, delayed rescue was associated with higher mortality, approximately 3.55 additional hospital days, and approximately $36,361 higher mean charges.

In the primary survey-weighted cross-fit AIPW analysis, early rescue remained associated with lower in-hospital mortality. The estimated mortality risk difference for early versus delayed rescue was −0.001556, standard error 0.000551, 95% CI −0.002636 to −0.000477. Expressed clinically, this corresponds to approximately 1.56 fewer in-hospital maternal deaths per 1,000 rescued high-risk delivery hospitalizations under early versus delayed rescue, conditional on the identification assumptions. Early rescue was also associated with shorter hospitalization: the AIPW mean difference in length of stay was −2.2267 days, standard error 0.0563, 95% CI −2.3371 to −2.1163. For total charges, the AIPW mean difference was −$15,933.68, standard error $878.45, 95% CI −$17,655.41 to −$14,211.96.

Diagnostics supported use of the weighted causal framework while identifying areas requiring careful interpretation. Propensity-score distributions showed substantial shared support between early and delayed rescue groups, although probabilities were concentrated toward early rescue, as expected given that 87.2% of rescued high-risk hospitalizations had rescue by hospital day 0–1. Covariate balance improved materially after stabilized weighting combined with discharge weights. The largest absolute standardized mean difference decreased from 0.212 before weighting to 0.055 after weighting, with most post-weighting standardized mean differences below the conventional 0.10 threshold. Missingness was minimal for the primary analysis: in-hospital mortality, length of stay, earliest rescue time, and rescue-timing category had 0% missingness in the analytic cohort; total charges had 0.76% missingness; age had 0.005% missingness.

Exploratory subgroup analyses suggested that the mortality contrast was most precise in larger strata. Among patients aged 30–39 years, early rescue was associated with a mortality risk difference of −0.00258, 95% CI −0.00448 to −0.00068, based on 27,193 unweighted hospitalizations. Among patients aged 20–29 years, the estimate was −0.00059, 95% CI −0.00178 to 0.00060, based on 23,296 hospitalizations. Among patients aged ≥40 years, the point estimate was −0.00248, but the interval was wide, 95% CI −0.00773 to 0.00277, based on 3,104 hospitalizations. By admission timing, weekday admissions showed a clearer mortality reduction, with risk difference −0.00175, 95% CI −0.00302 to −0.00048, based on 45,785 hospitalizations; weekend admissions had a smaller and less precise estimate, −0.00069, 95% CI −0.00277 to 0.00140, based on 10,178 hospitalizations. These subgroup estimates should be interpreted as exploratory because subgroup-specific confounding, outcome sparsity, and positivity may differ across strata.

The Obstetric Rescue Clock converts obstetric rescue delay into a national, procedure-day estimand for maternal safety. By treating first rescue day as the exposure, it estimates the survey-population contrast between early and delayed rescue among rescued high-risk deliveries, avoiding rescue-versus-nonrescue comparisons and improving causal comparability under confounding by indication. In NIS 2021, early rescue was associated with lower mortality, shorter stay, and lower charges. The framework links clinical timing, survey inference, doubly robust estimation, and bias-aware validation, making obstetric failure to rescue measurable, benchmarkable, and extensible to hierarchical surveillance. 

Keywords

Maternal mortality

Obstetric rescue timing

Survey-weighted inference

Doubly robust estimation

Target trial emulation

National Inpatient Sample (NIS) 

Speaker

Sunday Adetunji, Oregon State University

Sampling and Weighting Strategies for Nonprobability Survey and Serology Data from Blood Donors

Monitoring the burden of respiratory virus infections is essential to public health surveillance and preparedness for novel viral agents with pandemic potential. In the U.S., longitudinal survey and multipathogen serologic data are being collected from a cohort of blood donors under the Respiratory Virus Repeat Donor Cohort program, sponsored by the Centers for Disease Control and Prevention. Estimates of interest include incidence of respiratory virus infections nationally, regionally, and by demographic group. A sampling and weighting strategy was developed to reduce bias from nonprobability components while balancing representation and feasibility. Using donor lists from Vitalant and American Red Cross, an initial sample was selected via stratified random sampling calibrated to American Community Survey demographic totals. A series of probability sampling and non-sampling steps produced a cohort of 25,255 donors invited to complete quarterly surveys and a subset of 11,202 whose blood donations were tested for antibodies to multiple respiratory viruses. Weighting adjustments and calibration improved population representativeness, enhancing the data's utility for decision making. 

Keywords

Calibration

Nonprobability samples

Public health

Stratification

Subsampling 

Speaker

Laura Gamble, Westat

Co-Author(s)

Elizabeth Eisenhauer, Westat
Andreea Erciulescu, Westat
Rebecca Fink, Westat
Sarah Smith-Jeffcoat, Centers for Disease Control and Prevention
Jefferson Jones, Centers for Disease Control and Prevention
Bryan Spencer, American Red Cross
Eduard Grebe, Vitalant Research Institute
Mars Stone, Vitalant Research Institute
Michael Busch, Vitalant Research Institute

Markov Missing Graph: A Graphical Approach for Missing Data Imputation

We introduce the Markov missing graph (MMG), a novel framework that imputes missing data based on undirected graphs. MMG leverages conditional independence relationships to locally decompose the imputation model. To establish the identification, we introduce the Principle of Available Information (PAI), which guides the use of all relevant observed data. We then propose a flexible statistical learning paradigm, MMG Imputation Risk Minimization under PAI, that frames the imputation task as an empirical risk minimization problem. This framework is adaptable to various modeling choices. We develop theories of MMG, including the connection between MMG and Little's complete-case missing value assumption, recovery under missing completely at random, efficiency theory, and graph-related properties. We show the validity of our method with simulation studies and illustrate its application with a real-world Alzheimer's data set. 

Keywords

imputation

missing data

missing not at random

undirected graph

semi-parametric efficiency 

Speaker

Yanjiao Yang

Co-Author

Yen-Chi Chen, University of Washington

State space models for multivariate time series data with nonmonotone MNAR missingness

State-space models are widely used for analyzing multivariate time series with missing observations. In many longitudinal studies, the pattern of missingness is nonmonotone, and the probability of missingness depends directly on unobserved signal values, leading to nonmonotone Missing Not At Random (MNAR) mechanisms that induce systematic bias under the widely popular MAR-based state-space inferences developed in the literature so far. We propose linear-Gaussian state-space models with explicit MNAR modeling for such data, where the missingness process is governed by a logistic regression depending on latent states, current and lagged emissions, and historical missingness. By leveraging Pólya--Gamma augmentation, the MNAR mechanism is rendered conditionally conjugate. This enables closed-form posterior updates via Gibbs sampling and variational inference via Kalman filtering and Rauch--Tung--Striebel smoothing within a coordinate-ascent algorithm. Through simulation studies and an application on optical sensors for AQI in the Montana SmartFIRES study, we demonstrate the advantages of our method over existing state-of-the-art approaches. 

Keywords

State space models

nonmonotone MNAR

Kalman filtering and smoothing

Variational Inference 

Speaker

Kunal Das, Montana State University

Co-Author

Michael Wojnowicz, Montana State University

Variance Estimation for Local Pivotal Method via Monte Carlo Modeling of Joint Inclusion Probabilities

Spatially balanced designs, like the local pivotal method (LPM), reduce variance by inducing negative dependence among nearby units in covariate space. While the Horvitz-Thompson estimators remain design-unbiased, variance estimation is challenging because second-order inclusion probabilities are often intractable or near zero. Existing solutions follow two paths: Monte Carlo simulation to empirically estimate these probabilities, which is computationally prohibitive for large populations, or neighborhood-based estimators that bypass joint probabilities by sacrificing the use of standard design-unbiased estimators.
We link these approaches via a Monte Carlo modeling framework for variance estimation under LPM. A key limitation with Monte Carlo simulation is, without a computationally prohibitive number of samples, the empirical inclusion probabilities for spatially distant, independent pairs will falsely deviate from independence due to Monte Carlo noise. This noise inflates variance estimates which use purely empirical inclusion probabilities. To address this, we empirically estimate joint-inclusion probabilities for a subset of unit pairs, including all sampled pairs, using a feasible number of samples and train a logistic regression model to estimate the probability that a pair is independent given its covariate distance. Then, we calculate final predicted inclusion probabilities by blending the empirical estimates with the probability of independence. Using these predicted inclusion probabilities produces variance estimates that outperform existing methods without the computational burden of pure Monte Carlo estimation.  

Keywords

Local Pivotal Method

Spatially Balanced Designs

Variance Estimation

Second-Order Inclusion Probabilities

Monte Carlo 

Speaker

Justin Greene, Rutgers University

Co-Author

Tirthankar Dasgupta, Rutgers University

Considerations and Results on Fay-Train Variance Replication

Fay-Train replication guides many of the U.S. Census Bureau's variance estimation practices, particularly for the American Community Survey (ACS) and the Current Population Survey (CPS). The replication method begins by internally linearizing a positive semi-definite quadratic form used to represent the variance structure of a total estimator. This linearization produces a set of replication factors which, when differenced against survey weights and squared, recover the original quadratic form and, thus, the total's variance. This form of replication factor is subsequently used to construct variance estimators for non-linear functions, as well as both linear and non-linear variance estimators under raking and calibration constraints. This paper examines the vector and matrix forms of Fay-Train replication, clarifies its use and common misconceptions, discusses the properties of Fay-Train when the discrepancy between replication weights and survey weights are arbitrarily small, and briefly introduces related accumulated and ongoing research. 

Keywords

replication methods

survey variance estimation

Fay-Train linearization

official statistics

weight calibration

American Community Survey (ACS) 

Speaker

Patrick Joyce, United States Census Bureau