AI Approaches in Microbial Community Multi-omics

Place Holder Chair
 
Himel Mallick Organizer
Cornell University
 
Wednesday, Aug 5: 10:30 AM - 12:20 PM
1352 
Invited Paper Session 
Thomas M. Menino Convention & Exhibition Center 
Room: CC-210A 

Applied

Yes

Main Sponsor

Section on Statistics in Genomics and Genetics

Co Sponsors

Biometrics Section
Biopharmaceutical Section

Presentations

Multistudy Multimodal Pretraining and Transfer Learning

Machine learning has become integral to biomedical research, yet its success often depends on large, high-quality datasets that are rarely available. Transfer learning (TL) offers a principled way to leverage information from related large-scale datasets to improve prediction in a target study. However, existing TL methods are typically designed for single-view settings, limiting their use with complex multiview data collected across multiple cohorts. To address this gap, we propose a Multistudy Multimodal TL framework that enables integration-aware knowledge transfer across studies and modalities. The framework employs pretraining using cooperative learning followed by fine-tuning, yielding both general and context-specific predictions. Simulation studies and multimodal analyses from cancer immunotherapy and inflammatory bowel disease cohorts demonstrate that our method improves predictive accuracy and cross-study generalization. Our method thus establishes a principled and scalable foundation for cross-study and cross-modality transfer learning, substantially outperforming published methods in estimation and prediction. An open-source implementation is publicly available. 

Keywords

Transfer learning

Cooperative Learning

Multimodal Integration

Multistudy

Pretraining

Regularization 

Speaker

Chuxuan Gao, Weill Cornell Medicine

Zentangler: A Multimodal Mediation Analysis Framework for Multiview Data

Multimodal artificial intelligence (AI) has advanced considerably over the past decade, making multimodal data integration increasingly central to biomedical research. However, most existing analytic approaches remain limited to association-based, single-modality modeling, and there is currently no unified framework for causal mediation analysis across biological layers.

To address this gap, we introduce a population-scale multimodal mediation framework for characterizing causal relationships across multiomics data. The framework accommodates both parallel mediation structures (e.g., X → M₁ → Y and X → M₂ → Y) and sequential pathways (e.g., X → M₁ → M₂ → Y), enabling systematic decomposition of cross-modal interactions and their influence on downstream outcomes.

We present Zentangler, a unified multimodal AI engine that integrates early, intermediate, and late fusion strategies, and supports continuous, binary, survival, and multiclass outcomes. This engine is embedded within a counterfactual mediation framework, allowing decomposition of total effects into natural direct and indirect components across complex mediation pathways.

Across large-scale microbiome multiomics datasets, including iHMP and FINRISK, Zentangler outperforms existing single-modality approaches and reveals biologically meaningful cross-modal relationships. Synthetic benchmarking further demonstrates its improved accuracy and robustness. The method is implemented as an open-source R/Bioconductor package, available at: https://github.com/himelmallick/Zentangler/

 

Keywords

Multimodal AI

Causal Mediation Analysis

Multimodal Integration

Microbiome

Multiomics 

Speaker

Nalin Arora

Co-Author(s)

Nalin Arora
Himel Mallick, Cornell University

Inference on covariance structure in high-dimensional multi-view data

This work focuses on covariance estimation for multi-view data. Popular approaches rely on factor-analytic decompositions that have shared and view-specific latent factors. Posterior computation is conducted via expensive and brittle Markov chain Monte Carlo (MCMC) sampling or variational approximations that underestimate uncertainty and lack theoretical guarantees. Our proposed methodology employs spectral decompositions to estimate and align latent factors that are active in at least one view. Conditionally on these factors, we choose jointly conjugate prior distributions for factor loadings and residual variances. The resulting posterior is a simple product of normal-inverse gamma distributions for each variable, bypassing MCMC and facilitating posterior computation. We prove favorable increasing-dimension asymptotic properties, including posterior contraction and central limit theorems for point estimators. We show excellent performance in simulations, including accurate uncertainty quantification, and apply the methodology to integrate four high-dimensional views from a multi-omics dataset of cancer cell samples. 

Keywords

Factor analysis

High-dimensional

Latent variable model

Multi-omics

Scalable Bayesian computation

Singular value decomposition 

Speaker

Lorenzo Mauri

Tweedieverse and DAssemble: Robust Multimodal Analysis for Microbiome Data

The reproducibility crisis in omics data science is particularly evident in microbiome research, where heterogeneity in sequencing technologies, library preparation protocols, and preprocessing pipelines leads to complex count data characterized by sparsity, overdispersion, and compositional effects. Despite this variability, most differential abundance methods rely on a single distributional assumption across all features, resulting in model misspecification, unstable inference, and inconsistent findings. We present two complementary frameworks to mitigate this gap, Tweedieverse and DAssemble. Tweedieverse leverages the flexibility of the Tweedie distribution to enable feature-specific distributional adaptation through a single tuning parameter, providing a unified approach that captures diverse mean–variance relationships across features and modalities. DAssemble, on the other hand, improves robustness by aggregating differential abundance signals across multiple statistical models, thereby mitigating model-specific biases and reducing sensitivity to any single modeling assumption. Analyses of real and synthetic microbiome datasets demonstrate improved statistical power, false discovery rate control, and stability of discoveries relative to published methods. Open-source R packages implementing these methods are publicly available at https://github.com/himelmallick/DAssemble and https://github.com/himelmallick/tweedieverse. 

Keywords

Differential Analysis

Microbiome

Multimodal

Multi-omics

Metagenomics 

Speaker

Ziyu Liu

Co-Author(s)

Himel Mallick, Cornell University
Ziyu Liu