35: Multivariate Random Forest-Based Clustering for Integrative Multi-Omics Analysis
Xi Chen
Co-Author
University of Miami
Monday, Aug 3: 2:00 PM - 3:50 PM
3122
Contributed Posters
Thomas M. Menino Convention & Exhibition Center
Bulk multi-omics clustering methods either fuse modalities into a single similarity graph and lose modality-specific signal, or decompose variation into joint and individual factors under linear and distributional assumptions that rarely hold across layers with different marginal characteristics. No existing framework pairs a flexible, distribution-free similarity with an explicit shared and specific decomposition that can be clustered independently. We introduce multiRF, which addresses this gap using directed multivariate random forests fitted between all pairs of omics blocks. For each directed connection, a multivariate random forest is trained with one omics block as predictor and another as response, producing a sample-level weight matrix whose rows encode cross-modal predictive neighbourhoods. Because splits are axis-aligned and operate feature by feature, the forests accommodate mixed data types, nonlinear inter-omics relationships, and the high-dimensional, low-sample-size setting common to molecular profiling, all without requiring a parametric functional form or matched feature scales across modalities. A response subsampling mechanism further regulates the multivariate split criterion when the response block is high-dimensional, ensuring that the weight matrix reflects genuine cross-modal structure rather than spurious high-dimensional correlations. The per-connection weight matrices are then fused into a single global weight matrix that captures shared structure across all modalities. This matrix partitions every omics block into a shared reconstruction driven by cross-modal consensus and a residual that isolates modality-specific variation, without explicit latent-factor estimation or rank assumptions. Shared and specific similarity matrices are constructed from these two components and clustered independently, so that cross-omics disease subtypes and modality-private biological axes both emerge from a single pipeline with full traceability to the underlying inter-omics predictive relationships. multiRF is available as an open-source R package at https://github.com/novawz/multiRF.
Multivariate Random Forest
Multi-Omics Integration
High-Dimensional Data
Main Sponsor
Section on Statistical Learning and Data Science
You have unsaved changes.