Robust and Explainable Statistical Learning for Large-Scale Neuroimaging and Brain Networks

Manjun Yu Chair
 
Xiaojing Wang Organizer
University of Connecticut
 
Panpan Zhang Organizer
Vanderbilt University Medical Center
 
Tuesday, Aug 4: 10:30 AM - 12:20 PM
1793 
Topic-Contributed Paper Session 
Thomas M. Menino Convention & Exhibition Center 
Room: CC-253B 

Applied

Yes

Main Sponsor

Section on Statistics in Imaging

Co Sponsors

Section on Bayesian Statistical Science
Section on Statistical Learning and Data Science

Presentations

A Bayesian Model for Network Analysis under Edge Misclassification

Functional brain networks derived from resting-state fMRI are widely used to investigate connectivity alterations in biomedical research. However, such networks are often corrupted by measurement error: noisy correlations can introduce spurious edges (false positives), while low signal-to-noise ratio or infarcts may obscure true connections (false negatives). Standard network models typically ignore this uncertainty, leading to biased estimation of network structure and reduced reproducibility. We propose a latent space covariate model (LSCM) with explicit link misclassification to address these challenges. The model extends classical latent space approaches by incorporating node-level covariates and dyadic similarities while capturing hidden structure through latent positions. Misclassification is parameterized through both false positive and false negative error rates, yielding an observed network likelihood that is an affine transformation of the underlying true edge probabilities. Inference is carried out within a Bayesian framework, with prior distributions carefully chosen to regularize parameters that are otherwise weakly identifiable. Simulation studies demonstrate that explicitly modeling link misclassification improves recovery of latent positions, covariate effects, and global network properties compared to naive approaches. We further apply the proposed model to resting-state fMRI data in Alzheimer's disease, where it identifies robust and biologically meaningful subnetworks associated with cognitive decline. 

Keywords

network analysis

rs-fMRI data

erroneous links

latent space model

Alzheimer's disease 

Speaker

Panpan Zhang, Vanderbilt University Medical Center

Deep Generative Modeling with Spatial and Network Images: An Explainable AI (XAI) Approach

This talk will address the challenge of modeling the amplitude of spatially indexed low frequency fluctuations (ALFF) in resting state functional MRI as a function of cortical structural features and a multi-task coactivation network in the Adolescent Brain Cognitive Development (ABCD) Study. It proposes a generative model that integrates effects of spatially-varying inputs and a network-valued input using deep neural networks to capture complex non-linear and spatial associations with the output. The method models spatial smoothness, accounts for subject heterogeneity and complex associations between network and spatial images at different scales, enables accurate inference of each images effect on the output image, and allows prediction with uncertainty quantification via Monte Carlo dropout, contributing to one of the first Explainable AI (XAI) frameworks for heterogeneous imaging data. The model is highly scalable to high-resolution data without the heavy pre-processing or summarization often required by Bayesian methods. Empirical results demonstrate its strong performance compared to existing statistical and deep learning methods. We applied the XAI model to the ABCD data which revealed associations between cortical features and ALFF throughout the entire brain. Our model performed comparably to existing methods in predictive accuracy but provided superior uncertainty quantification and faster computation, demonstrating its effectiveness for large-scale neuroimaging analysis. Open-source software in Python for XAI is available. 

Speaker

Rajarshi Guhaniyogi, Texas A&M University

Improved Estimation of Correlation Accuracy for Machine Learning Brain-Phenotype AssociationsPresentation

Machine learning is used in neuroscience to examine brain-phenotype associations and facilitate individual prediction from high-dimensional brain imaging. For continuous phenotypes, Pearson's correlation between the observed and predicted phenotype is used to quantify model accuracy in testing data. However, recent research suggests millions of samples may be needed to reliably estimate the maximum achievable predictive accuracy (MAPA). We formally define the MAPA and show that Pearson's estimator is biased for this quantity and its confidence intervals fail to capture the target. We develop a semiparametric (double machine learning) one-step estimator that more accurately estimates the MAPA and yields valid confidence intervals across flexible machine learning settings. Analyzing data from the Reproducible Brain Charts dataset, we show that this estimator has smaller bias when estimating brain-phenotype associations of neuroimaging data with age and psychopathology phenotypes. We show that MAPA for psychopathology factor scores using machine learning models built on structural and functional imaging measures is not better than using demographic and nuisance covariates alone. 

Speaker

Simon Vandekar, Vanderbilt University Medical Center

Co-Author(s)

Megan Jones
Ishaan Gadiyar, Vanderbilt University Medical Center
Xinyu Zhang
Kaidi Kang, Wake Forest University School of Medicine
Jinyuan Liu, Vanderbilt University
Andrew Chen
Edward Kennedy

Longitudinal Mixed Membership Image-on-Scalar Model

Magnetic resonance imaging (MRI) data has been extensively applied in diagnosing
and predicting Alzheimer's disease (AD). However, there has been a notable oversight
in addressing individual heterogeneity within longitudinal MRI data. This paper introduces
a novel modeling framework to elucidate the diverse dynamic patterns inherent
in longitudinal imaging data, thereby facilitating a better understanding of individualized
AD progression. The framework commences with a basis expansion approach
to approximate the longitudinal images. Subsequently, a vector of probability weights
is introduced, delineating a subject's partial membership across clusters. Such partial
membership allows the subject's repeatedly measured imaging data to belong to different
clusters. Finally, a nonlinear trajectory model is employed to capture the typical
normal aging process and the potential transition from normal to severe stages during
the disease course. A Bayesian approach coupled with efficient MCMC algorithms is
developed for statistical inference. Extensive simulation results demonstrate the efficacy
of the proposed methods in parameter estimation and model selection. The framework
is applied to the Alzheimer's Disease Neuroimaging Initiative (ADNI) dataset, yielding
insights into the evolving patterns of brain structures throughout AD progression. 

Keywords

Mixed membership model

Imaging data

Longitudinal analysis

Alzheimer's disease 

Speaker

Xinyuan Song, The Chinese University of Hong Kong

Robust normality transformation for outlier detection in diverse distributions, with application to functional neuroimaging data

Automatic detection of statistical outliers is facilitated through knowledge of the source distribution of regular observations. Since the population distribution is often unknown in practice, one approach is to apply a transformation to Normality. However, the efficacy of transformation is hindered by the presence of outliers, which can have an outsized influence on transformation parameter(s) and lead to masking of outliers post-transformation. Robust Box-Cox and Yeo-Johnson transformations have been proposed but those transformations are only equipped to deal with skew. Here, we develop a novel robust method for transformation to Normality based on the highly flexible sinh-arcsinh (SHASH) family of distributions, which can accommodate skew, non-Gaussian tail weights, and combinations of both. A critical step is initializing outliers, given their potential influence on the highly flexible SHASH transformation. To this end, we consider conventional robust z-scoring and a novel anomaly detection approach. Through extensive simulation studies and real data analyses representing a wide variety of distribution shapes, we find that SHASH transformation outperforms existing methods, exhibiting high sensitivity to outliers even at heavy contamination levels (20-30\%). We illustrate the utility of SHASH transformation-based outlier detection in the context of noise reduction in functional neuroimaging data. 

Speaker

Saranjeet Singh Saluja

Co-Author(s)

Saranjeet Singh Saluja
Fatma Parlak, Indiana University
Amanda Mejia, Indiana University