Wednesday, Aug 5: 8:30 AM - 10:20 AM
6444
Contributed Papers
Thomas M. Menino Convention & Exhibition Center
Room: CC-257A
Main Sponsor
Survey Research Methods Section
Presentations
Prior research on childhood vaccination coverage surveys in Nigeria has shown that survey-to-survey changes in national outcome estimates can be notably amplified by changes in relative sums of stratum-level survey weights. To learn whether this is common or rare we analyzed sequences of Demographic and Health Surveys (DHS) and Multiple Indicator Cluster Surveys (MICS) from 29 African countries over three decades. We compared reported changes in national coverage estimates with counterfactual changes computed using fixed subnational weights. The contribution of weights was usually quite small; it was most pronounced in Nigeria, where it was evident in sequences of DHS and MICS surveys. This presentation describes the analytic approach, multi-country results, and recommendations for what to report to stakeholders about this issue.
Keywords
survey analysis
childhood vaccination coverage
DHS and MICS surveys
DPT3 and Meseales surveys
post-stratification
influence of survey weights
The growing demand for open data has intensified the need for release
strategies that protect respondent confidentiality while preserving the
validity of statistical analyses. Partially synthetic data, which
selectively replaces high-risk records with model-generated values, offers
a promising solution. However, when minority group members are
disproportionately flagged as high risk and targeted for synthesis, standard
generative models may distort subgroup-specific statistics, raising fairness
concerns that existing machine learning approaches do not adequately address
within traditional survey statistical frameworks.
We propose two Bayesian log-linear modeling approaches for balancing
fairness and confidentiality in partially synthetic data generation. The
first modifies the prior distribution to penalize deviations between
synthetic and original subgroup outcome rates. The second incorporates
fairness information through fixed offsets in the synthesis model. Both
methods are designed to preserve subgroup-level statistics without
excessively compromising privacy protection. We evaluate the methods using
2022 American Community Survey data from 12 Midwestern states, targeting
health insurance coverage rates across gender, race/ethnicity, and education
subgroups.
Results show that both methods substantially reduce subgroup-level fairness
error relative to the standard baseline, particularly in larger states with
greater sample sizes. Disclosure risk assessment confirms that fairness
improvements can be achieved without materially increasing re-identification
risk. A systematic fairness-privacy trade-off analysis using a weighted
Composite criterion reveals that the relationship between tuning parameters
and optimal performance is state-specific, with the prior-based method
offering a more stable trade-off across diverse settings. Overall, we
provide a generalizable framework for calibrating fairness and privacy in
synthetic data release and offer practical guidance for statistical agencies
seeking to produce datasets that are simultaneously secure, analytically
useful, and fair across diverse population subgroups.
Keywords
Disclosure limitation
Partially synthetic data
Fairness constraints
Privacy–utility trade-off
American Community Survey
Probability surveys face rising non-response rates, resulting in biased statistical inference. Auxiliary information can be used to reduce bias in estimation. Continuous auxiliary variables in administrative data are often discretized prior to release to avoid confidentiality breaches. This may limit the utility of administrative records in improving survey estimates, especially when continuous auxiliary data strongly predict the survey outcome. We propose a two-step strategy. First, statistical agencies use confidential continuous auxiliary data to estimate response propensity score of the survey sample and include these in a modified population data. Data users then conduct predictive inference including the discretized continuous variables and the propensity scores as predictors using splines in a Bayesian model. The proposed method performs well, yielding more efficient estimates of population means with 95% credible intervals providing better coverage than alternative approaches. We illustrate the proposed method using the Ohio Army National Guard Mental Health Initiative. The methods developed in this work are readily available in the R package AuxSurvey.
Keywords
Bayesian generalized additive model
continuous auxiliary variables
data protection
Rstan
response propensity
poststratification
We take advantage of a natural experiment implemented at October 2022 around policy intervention which allows merchants to surcharge for credit card transaction. We present empirical evidence on merchant surcharge and payment acceptance decisions measured by indirect sampling approach. We employ the data integration approach for the indirectly sampled data under the bipartite graph (i.e., consumer-merchant network linked by transactions). Compared to the existing weight share method, our proposed weights do neither require knowing the inclusion probabilities of directly sampled units, nor successors' and ancestors' knowledge of the bipartite graph. In a simulation, we assess the effectiveness of our weights in terms of bias and variance.
Keywords
Indirect sampling
Non-probability survey
Data integration
Car seats are designed to ensure that children are traveling safely in vehicles, but are frequently misused or not used at all, compromising child passenger safety. No nationally representative study provides recent estimates of car seat misuse. The National Digital Car Seat Check Form (NDCF) is an administrative data source that provides a national database of over 300,000 car seat checks performed by licensed child passenger safety technicians, including detailed information on car seat use and misuse, but without any personally identifying information other than zip code. However, it only covers children who attend car seat checks recorded in the NDCF.
By attaching ACS 5-year estimates via county, we can estimate the proportion of families with children under age 8 in a given area that are covered by the NDCF. We also examine other county-level characteristics, such as demographics and commuting behaviors, to assess whether any of these characteristics are associated with car seat check participation or car seat misuse rates. These variables are then used to develop weights for adjusted estimates that better represent misuse rates for all child passengers nationally.
Keywords
administrative data
weighting
child passenger safety
car seats
The Rapid Surveys System (RSS), managed by the National Center for Health Statistics (NCHS), is used to produce relevant and timely estimates on emerging health topics. Unlike other NCHS health surveys that use direct recruitment, the RSS surveys are administered through commercial online panels.
Previous work has shown differences between some RSS and National Health Interview Survey (NHIS) estimates. Statistical matching is one approach that can be used to remediate such differences by using a high-quality reference survey, such as the NHIS, as a basis. We consider two echelons of matching variables based on public-use data: an initial set consisting of socio-demographic characteristics that require exact matches between donors and recipients, and a second set of matching variables comprising health outcomes which will be used to generate competing distance matrices. This dual system can replace a large matrix with a more manageable collection of lower-dimensional matrices. These distances are not used deterministically but will inform probabilities of selection under repeated matchings. Monte Carlo estimates are then compared to NHIS variables not used in the matching proces
Keywords
Statistical Matching
Survey Statistics
Imputation
Monte Carlo
Bipartite record linkage is the problem of merging two duplicate-free databases in the absence of unique identifiers. Bayesian approaches to this problem offer natural uncertainty quantification and transitivity of matching decisions. However, these approaches rely on Markov chain Monte Carlo (MCMC) for posterior inference, limiting their scalability to small sized databases. In this talk, we propose a variational approximation for a Bayesian bipartite record linkage model. We use hashing and a re-parameterization of the approximating variational distribution to derive a coordinate ascent algorithm with complexity that grows linearly with the number of records in the smaller database. Additionally, we derive a stochastic variational inference algorithm with complexity that is independent of the database sizes. Through a series of simulations and applications, we illustrate that the variational approximations attain comparable accuracy to MCMC based methods, at a significantly decreased computational cost. Specifically, after pre-processing, we are able to merge 400,000 records in under 5 minutes, and 40,000 records in under 5 seconds.
Keywords
entity resolution
record linkage
variational inference
Bayesian methods