Wednesday, Aug 5: 10:30 AM - 12:20 PM
1352
Invited Paper Session
Thomas M. Menino Convention & Exhibition Center
Room: CC-210A
Applied
Yes
Main Sponsor
Section on Statistics in Genomics and Genetics
Co Sponsors
Biometrics Section
Biopharmaceutical Section
Presentations
Machine learning has become integral to biomedical research, yet its success often depends on large, high-quality datasets that are rarely available. Transfer learning (TL) offers a principled way to leverage information from related large-scale datasets to improve prediction in a target study. However, existing TL methods are typically designed for single-view settings, limiting their use with complex multiview data collected across multiple cohorts. To address this gap, we propose a Multistudy Multimodal TL framework that enables integration-aware knowledge transfer across studies and modalities. The framework employs pretraining using cooperative learning followed by fine-tuning, yielding both general and context-specific predictions. Simulation studies and multimodal analyses from cancer immunotherapy and inflammatory bowel disease cohorts demonstrate that our method improves predictive accuracy and cross-study generalization. Our method thus establishes a principled and scalable foundation for cross-study and cross-modality transfer learning, substantially outperforming published methods in estimation and prediction. An open-source implementation is publicly available.
Keywords
Transfer learning
Cooperative Learning
Multimodal Integration
Multistudy
Pretraining
Regularization
Multimodal artificial intelligence (AI) has advanced considerably over the past decade, making multimodal data integration increasingly central to biomedical research. However, most existing analytic approaches remain limited to association-based, single-modality modeling, and there is currently no unified framework for causal mediation analysis across biological layers.
To address this gap, we introduce a population-scale multimodal mediation framework for characterizing causal relationships across multiomics data. The framework accommodates both parallel mediation structures (e.g., X → M₁ → Y and X → M₂ → Y) and sequential pathways (e.g., X → M₁ → M₂ → Y), enabling systematic decomposition of cross-modal interactions and their influence on downstream outcomes.
We present Zentangler, a unified multimodal AI engine that integrates early, intermediate, and late fusion strategies, and supports continuous, binary, survival, and multiclass outcomes. This engine is embedded within a counterfactual mediation framework, allowing decomposition of total effects into natural direct and indirect components across complex mediation pathways.
Across large-scale microbiome multiomics datasets, including iHMP and FINRISK, Zentangler outperforms existing single-modality approaches and reveals biologically meaningful cross-modal relationships. Synthetic benchmarking further demonstrates its improved accuracy and robustness. The method is implemented as an open-source R/Bioconductor package, available at: https://github.com/himelmallick/Zentangler/
Keywords
Multimodal AI
Causal Mediation Analysis
Multimodal Integration
Microbiome
Multiomics
This work focuses on covariance estimation for multi-view data. Popular approaches rely on factor-analytic decompositions that have shared and view-specific latent factors. Posterior computation is conducted via expensive and brittle Markov chain Monte Carlo (MCMC) sampling or variational approximations that underestimate uncertainty and lack theoretical guarantees. Our proposed methodology employs spectral decompositions to estimate and align latent factors that are active in at least one view. Conditionally on these factors, we choose jointly conjugate prior distributions for factor loadings and residual variances. The resulting posterior is a simple product of normal-inverse gamma distributions for each variable, bypassing MCMC and facilitating posterior computation. We prove favorable increasing-dimension asymptotic properties, including posterior contraction and central limit theorems for point estimators. We show excellent performance in simulations, including accurate uncertainty quantification, and apply the methodology to integrate four high-dimensional views from a multi-omics dataset of cancer cell samples.
Keywords
Factor analysis
High-dimensional
Latent variable model
Multi-omics
Scalable Bayesian computation
Singular value decomposition
The reproducibility crisis in omics data science is particularly evident in microbiome research, where heterogeneity in sequencing technologies, library preparation protocols, and preprocessing pipelines leads to complex count data characterized by sparsity, overdispersion, and compositional effects. Despite this variability, most differential abundance methods rely on a single distributional assumption across all features, resulting in model misspecification, unstable inference, and inconsistent findings. We present two complementary frameworks to mitigate this gap, Tweedieverse and DAssemble. Tweedieverse leverages the flexibility of the Tweedie distribution to enable feature-specific distributional adaptation through a single tuning parameter, providing a unified approach that captures diverse mean–variance relationships across features and modalities. DAssemble, on the other hand, improves robustness by aggregating differential abundance signals across multiple statistical models, thereby mitigating model-specific biases and reducing sensitivity to any single modeling assumption. Analyses of real and synthetic microbiome datasets demonstrate improved statistical power, false discovery rate control, and stability of discoveries relative to published methods. Open-source R packages implementing these methods are publicly available at https://github.com/himelmallick/DAssemble and https://github.com/himelmallick/tweedieverse.
Keywords
Differential Analysis
Microbiome
Multimodal
Multi-omics
Metagenomics