Multiscale Data Integration Techniques for Official Statistics and Health

Sanjay Chaudhuri Chair
University of Nebraska-Lincoln
 
Sanjay Chaudhuri Organizer
University of Nebraska-Lincoln
 
Thursday, Aug 6: 8:30 AM - 10:20 AM
1723 
Topic-Contributed Paper Session 
Thomas M. Menino Convention & Exhibition Center 
Room: CC-258A 

Applied

Yes

Main Sponsor

Survey Research Methods Section

Co Sponsors

Government Statistics Section
Social Statistics Section

Presentations

Inference and prediction in multiscale data models

We consider statistically principled inference and prediction in some models that are routinely used in the context of multiscale data. The first part of our studies are for multiscale Gaussian processes, where we establish Bayesian statistical properties leading to inferences and predictions with probabilistic guarantees. Then we discuss some artificial neural network based models for multiscale data, and compare the results from such models to those using multiscale Gaussian process models. Numerical examples in the context of small area models and neuroimaging data are presented.
 

Speaker

Snigdhansu Chatterjee, University of Maryland, Baltimore County

Co-Author

Snigdhansu Chatterjee, University of Maryland, Baltimore County

A Bayesian Approach to Produce Subnational Population Estimates Using a Population Base Statistical Register

Subnational Population Estimates (SPE) in Latin America are useful to implement new public policies in subnational areas with internal armed conflicts or difficult to access. In this work, we propose to combine a Population Base Statistical Register (PBSR) and the Official Population Projections (OPP) using a Bayesian approach to produce SPE. Our proposed procedures are useful for computing SPE of the population size or the SPE of the population size in percentage SPE (%). However, we focused on SPE (%) due to some data restrictions and to ensure data confidentiality. In this article, the PBSR is constructed using multiple administrative sources with registers from the health, education, vital statistics systems, tax registration, and, more importantly, the registers of the victims of the current internal armed conflict in Colombia. We also propose new fast Markov chain Monte Carlo algorithms to produce SPE (%) using data augmentation procedures to address the complications caused by the resulting joint posterior containing gamma functions. We implement our proposal to compute SPE (%) by age and sex groups in the municipality of Jamundí in Colombia which is currently affected by poverty, forced displacement, and the internal armed conflict and evaluate the accuracy with a Population Census. 

Keywords

Subnational Population Estimates

Bayesian procedures

Population Base Statistical Register

Official Population Projection

Pólya Inverse-Gamma Schemes

Official Statistics 

Speaker

Jairo Alberto Fuquene Patino

Multiple frame surveys in modern data integration

Multiple frame surveys provide effective ways to integrate multiple data sets from heterogeneous
sources. Well-motivated by traditional survey sampling, the scope of this design is unfortunately
too limited to address modern applications in data integration. The first part of the talk develops
methods for hypothesis testing often overlooked by survey sampling when data sets are obtained
from independent surveys. Because parameter of interest is not a finite population parameter,
we adopt the super population framework. This additional randomness introduces multitude
of dependence within and across multiple data sets through potential duplication and finite
population sampling so that quantifying uncertainty is much harder and challenging than in the
finite population framework. With this complication, the distributions of the inverse probability
weighted version of pivotal quantities in the i.i.d. setting becomes no longer parameter-free.
Our proposed methodology first develops asymptotic theory for multiple frame surveys with a
super population and then estimates complex parameter-dependent null distributions through
simulation and/or bootstrap. Our methods are illustrated with data analysis of the Wilms tumor
study. The second part is our attempt to apply the framework of multiple frame surveys to non-
probability samples. In the non-probability sample research, one combines a reference survey
and a non-probability sample whose missingness mechanism is unknown. We extend our
asymptotic theory to combining data from sample surveys and missing data. After reviewing
methodology in non-probability samples, we discuss limitations of this approach and potential
solutions using techniques from survey sampling such as record linkage 

Keywords

Multiple Frame Surveys

Data Integration 

Speaker

Takumi Saegusa, University of Maryland

Co-Author

Takumi Saegusa, University of Maryland

Local Polynomial Regression for Learning Dynamical Systems

We propose a local polynomial regression framework for discovering governing equations of dynamical systems from noisy time-indexed data. The method estimates both the latent state and its time derivative under Gaussian measurement error, illustrated using an exponential growth model. Closed-form local estimators enable recovery of the underlying dynamics without global parametric assumptions. The approach is extended to incorporate random effects, allowing for subject-specific heterogeneity in growth dynamics. Simulation studies demonstrate accurate recovery of the governing equations under moderate noise, highlighting the method's flexibility for longitudinal and repeated-measures data. 

Keywords

Dynamical systems

Local Polynomial Regression 

Speaker

Siddhartha Nandy, Case Western Reserve University

Empircal Likelihood-based Methods for Multiscale Data Integration

It is well-known that the empirical likelihood-based methods provide an easy way of data integration in a vast array of real-life problems. In this talk, we will discuss the application of such methods to multiscale data integration. Basic theory of empirical likelihood-based data integration will be introduced. Based on these preliminaries, we will discuss new methods for the integration of multiscale data. The methods would be illustrated with real-life examples. 

Keywords

Multiscale data

Empirical Likelihood

Bayesian empirical Likelihood

Approximate Bayesian Computation 

Speaker

Sanjay Chaudhuri, University of Nebraska-Lincoln