Data integration, disclosure limitation, and combined data approaches

Wednesday, Aug 5: 8:30 AM - 10:20 AM
6444 
Contributed Papers 
Thomas M. Menino Convention & Exhibition Center 
Room: CC-257A 

Main Sponsor

Survey Research Methods Section

Presentations

The Quiet Influence of Stratum-Level Survey Weights on Perceptions of Childhood Vaccination Coverage

Prior research on childhood vaccination coverage surveys in Nigeria has shown that survey-to-survey changes in national outcome estimates can be notably amplified by changes in relative sums of stratum-level survey weights. To learn whether this is common or rare we analyzed sequences of Demographic and Health Surveys (DHS) and Multiple Indicator Cluster Surveys (MICS) from 29 African countries over three decades. We compared reported changes in national coverage estimates with counterfactual changes computed using fixed subnational weights. The contribution of weights was usually quite small; it was most pronounced in Nigeria, where it was evident in sequences of DHS and MICS surveys. This presentation describes the analytic approach, multi-country results, and recommendations for what to report to stakeholders about this issue. 

Keywords

survey analysis

childhood vaccination coverage

DHS and MICS surveys

DPT3 and Meseales surveys

post-stratification

influence of survey weights 

Speaker

Neelakshi Chatterjee

Co-Author(s)

Caitlin Clary, Biostat Global Consulting
Dale Rhoda, Biostat Global Consulting LLC

Balancing Fairness and Disclosure Limitation when Creating Synthetic Data for Survey Research

The growing demand for open data has intensified the need for release
strategies that protect respondent confidentiality while preserving the
validity of statistical analyses. Partially synthetic data, which
selectively replaces high-risk records with model-generated values, offers
a promising solution. However, when minority group members are
disproportionately flagged as high risk and targeted for synthesis, standard
generative models may distort subgroup-specific statistics, raising fairness
concerns that existing machine learning approaches do not adequately address
within traditional survey statistical frameworks.

We propose two Bayesian log-linear modeling approaches for balancing
fairness and confidentiality in partially synthetic data generation. The
first modifies the prior distribution to penalize deviations between
synthetic and original subgroup outcome rates. The second incorporates
fairness information through fixed offsets in the synthesis model. Both
methods are designed to preserve subgroup-level statistics without
excessively compromising privacy protection. We evaluate the methods using
2022 American Community Survey data from 12 Midwestern states, targeting
health insurance coverage rates across gender, race/ethnicity, and education
subgroups.

Results show that both methods substantially reduce subgroup-level fairness
error relative to the standard baseline, particularly in larger states with
greater sample sizes. Disclosure risk assessment confirms that fairness
improvements can be achieved without materially increasing re-identification
risk. A systematic fairness-privacy trade-off analysis using a weighted
Composite criterion reveals that the relationship between tuning parameters
and optimal performance is state-specific, with the prior-based method
offering a more stable trade-off across diverse settings. Overall, we
provide a generalizable framework for calibrating fairness and privacy in
synthetic data release and offer practical guidance for statistical agencies
seeking to produce datasets that are simultaneously secure, analytically
useful, and fair across diverse population subgroups. 

Keywords

Disclosure limitation

Partially synthetic data

Fairness constraints

Privacy–utility trade-off

American Community Survey 

Speaker

Chendi Zhao, University of Michigan, Ann Arbor

Co-Author(s)

Trivellore Raghunathan, University of Michigan
Joerg Drechsler, Institute for Employment Research

Maximizing continuous auxiliary data for improved survey inference while controlling disclosure risk

Probability surveys face rising non-response rates, resulting in biased statistical inference. Auxiliary information can be used to reduce bias in estimation. Continuous auxiliary variables in administrative data are often discretized prior to release to avoid confidentiality breaches. This may limit the utility of administrative records in improving survey estimates, especially when continuous auxiliary data strongly predict the survey outcome. We propose a two-step strategy. First, statistical agencies use confidential continuous auxiliary data to estimate response propensity score of the survey sample and include these in a modified population data. Data users then conduct predictive inference including the discretized continuous variables and the propensity scores as predictors using splines in a Bayesian model. The proposed method performs well, yielding more efficient estimates of population means with 95% credible intervals providing better coverage than alternative approaches. We illustrate the proposed method using the Ohio Army National Guard Mental Health Initiative. The methods developed in this work are readily available in the R package AuxSurvey. 

Keywords

Bayesian generalized additive model

continuous auxiliary variables

data protection

Rstan

response propensity

poststratification 

Speaker

Sharifa Williams

Co-Author(s)

Jungang Zou
Yutao Liu
Yajuan Si, University of Michigan
Sandro Galea, Washington University School of Public Health
Qixuan Chen, Columbia University

Estimate the prevalence of merchant payment steering from transaction receipt data

We take advantage of a natural experiment implemented at October 2022 around policy intervention which allows merchants to surcharge for credit card transaction. We present empirical evidence on merchant surcharge and payment acceptance decisions measured by indirect sampling approach. We employ the data integration approach for the indirectly sampled data under the bipartite graph (i.e., consumer-merchant network linked by transactions). Compared to the existing weight share method, our proposed weights do neither require knowing the inclusion probabilities of directly sampled units, nor successors' and ancestors' knowledge of the bipartite graph. In a simulation, we assess the effectiveness of our weights in terms of bias and variance. 

Keywords

Indirect sampling

Non-probability survey

Data integration 

Speaker

Heng Chen, BoC

Estimating Car Seat Misuse Rates using Administrative Child Passenger Safety Data

Car seats are designed to ensure that children are traveling safely in vehicles, but are frequently misused or not used at all, compromising child passenger safety. No nationally representative study provides recent estimates of car seat misuse. The National Digital Car Seat Check Form (NDCF) is an administrative data source that provides a national database of over 300,000 car seat checks performed by licensed child passenger safety technicians, including detailed information on car seat use and misuse, but without any personally identifying information other than zip code. However, it only covers children who attend car seat checks recorded in the NDCF.

By attaching ACS 5-year estimates via county, we can estimate the proportion of families with children under age 8 in a given area that are covered by the NDCF. We also examine other county-level characteristics, such as demographics and commuting behaviors, to assess whether any of these characteristics are associated with car seat check participation or car seat misuse rates. These variables are then used to develop weights for adjusted estimates that better represent misuse rates for all child passengers nationally. 

Keywords

administrative data

weighting

child passenger safety

car seats 

Speaker

Elizabeth Petraglia, Westat

Statistical Matching of the Rapid Surveys System to the National Health Interview Survey

The Rapid Surveys System (RSS), managed by the National Center for Health Statistics (NCHS), is used to produce relevant and timely estimates on emerging health topics. Unlike other NCHS health surveys that use direct recruitment, the RSS surveys are administered through commercial online panels.

Previous work has shown differences between some RSS and National Health Interview Survey (NHIS) estimates. Statistical matching is one approach that can be used to remediate such differences by using a high-quality reference survey, such as the NHIS, as a basis. We consider two echelons of matching variables based on public-use data: an initial set consisting of socio-demographic characteristics that require exact matches between donors and recipients, and a second set of matching variables comprising health outcomes which will be used to generate competing distance matrices. This dual system can replace a large matrix with a more manageable collection of lower-dimensional matrices. These distances are not used deterministically but will inform probabilities of selection under repeated matchings. Monte Carlo estimates are then compared to NHIS variables not used in the matching proces 

Keywords

Statistical Matching

Survey Statistics

Imputation

Monte Carlo 

Speaker

William Waldron, National Center for Health Statistics

Co-Author(s)

Katherine Irimata, National Center for Health Statistics
Li-Yen Hu
James Dahlhamer, National Center for Health Statistics

Variational Beta Linkage

Bipartite record linkage is the problem of merging two duplicate-free databases in the absence of unique identifiers. Bayesian approaches to this problem offer natural uncertainty quantification and transitivity of matching decisions. However, these approaches rely on Markov chain Monte Carlo (MCMC) for posterior inference, limiting their scalability to small sized databases. In this talk, we propose a variational approximation for a Bayesian bipartite record linkage model. We use hashing and a re-parameterization of the approximating variational distribution to derive a coordinate ascent algorithm with complexity that grows linearly with the number of records in the smaller database. Additionally, we derive a stochastic variational inference algorithm with complexity that is independent of the database sizes. Through a series of simulations and applications, we illustrate that the variational approximations attain comparable accuracy to MCMC based methods, at a significantly decreased computational cost. Specifically, after pre-processing, we are able to merge 400,000 records in under 5 minutes, and 40,000 records in under 5 seconds. 

Keywords

entity resolution

record linkage

variational inference

Bayesian methods 

Speaker

Serge Aleshin-Guendel, United States Census Bureau

Co-Author(s)

Brian Kundinger, Duke University
Yinyihong Liu, Duke University
Rebecca Steorts, Duke University