Tuesday, Aug 4: 2:00 PM - 3:50 PM
6427
Contributed Papers
Thomas M. Menino Convention & Exhibition Center
Room: CC-204A
Main Sponsor
Section on Statistics in Genomics and Genetics
Presentations
In this talk, a Bayesian framework is proposed to cluster observations sharing a common undirected network. The clustering, denoted as Multi-center Graph Clustering (McGC), is driven by network structures and strength of edges and allows the variables in each cluster to have different profiles. Pseudo nodes designed with no expected connections with the original nodes are introduced to control false connections aiming to facilitate graph constructions. Extensive simulations demonstrate the feasibility of the proposed approach and applications of McGC to epigenetic data support its value in practice with a potential to benefit future studies in predicting disease risk at a much earlier stage of life.
Keywords
Gaussian graphs
Bayesian inference
Pseudo nodes
Cluster analysis
Tuning parameter
Variable selection
Human diseases have distinct manifestations across biological tissues. RNA expression within the cells of each tissue provides unique insight into disease progression. Commonly, only a single tissue can be measured due to factors such as cost or ease of sample collection. For example, it is more difficult to collect samples from internal organs, like the lungs and heart, than to obtain blood samples. We propose a method to predict RNA expression in an unmeasured tissue using RNA expression from a measured tissue. We draw inspiration from sparse reduced rank regression by Chen and Huang (2012), which uses a group lasso constraint to remove entire genes from the predictor set. We explore alternative sparsity constraints, such as the standard lasso, which allows predictor gene sets to vary depending on the response gene, and the exclusive lasso, in which each predictor may contribute to at most one latent factor. As we consider solutions to this problem, we navigate the trade-off between complexity and interpretability.
Keywords
Reduced rank regression (RRR)
Cross-tissue RNA-seq prediction
Transcriptomics
Lasso - standard, group, and exclusvie
Loss functions for counts data
Latent variable models
Modern multi-omics technologies profile multiple molecular layers (e.g., genomic, transcriptomic, metabolomic) to elucidate biological mechanisms. Low-rank factorization effectively reveals coordinated signatures across these layers representing shared physiological or pathological signals. However, most existing models are purely data driven, prioritizing predictive power over biological priors such as annotated pathways or gene sets. We propose a novel framework leveraging this knowledge to guide the low-rank factorization of multi-omics data. Unlike standard methods that seek generic latent structures, our algorithm constrains factors to represent distinct up- and down-regulated functional sets. This approach allows for a mechanistic interpretation where each latent factor corresponds to a specific regulatory state across omics layers. Extensive simulations confirm improved factor specialization and recovery of signed pathways and supporting features. Benchmarking on multi-omic cancer cell line profiles (DepMap), we show that our knowledge-centric approach improves stability and yields superior biological insight into regulatory heterogeneity compared to leading competitors.
Keywords
Multi-omics Integration
Matrix Factorization
Structured Penalization
Directional False Discovery Control
Pathway Analysis
Statistical Software
Canonical Correlation Analysis (CCA) involves identifying the linear combinations of two sets of measurements from the same set of observations that have the highest correlation. This method has been used in genomics data to explore associations between different data modalities, such as DNA copy number and single nucleotide polymorphisms (SNPs). In this study, we focus on differential CCA (dCCA), an extension of CCA that maximizes the absolute difference between canonical correlations from two distinct groups. This identifies linear combinations that maximally contrast the between-measurement associations in one group versus the other. We propose algorithms to solve both the sparse and dense linear combinations of variables. In addition, we propose a generalization of dCCA for multiple groups. This method finds linear combinations such that the resulting canonical correlation from each group exhibits the strongest possible association with a continuous outcome variable (e.g., a risk score) assigned to that group. Finally, we explore applications of dCCA in studying cell-cell interactions as well as axial patterns in spatial transcriptomics data.
Keywords
spatial transcriptomics
canonical correlation analysis
Speaker
Yuzi Li, Johns Hopkins University, Department of Biostatistics
Co-Author
Hongkai Ji, Johns Hopkins University
Data-Independent Acquisition (DIA) proteomics generates precursor (MS1) and fragment (MS2) ion data. Despite their value, different signal interferences cause inconsistencies when analyzed separately. Hypothesizing that integrating MS1 and MS2 signals improves differential protein abundance analysis accuracy, we developed a linear mixed-effects model (LMM) that jointly analyzes MS1 and MS2 intensities-normalized to exosome markers-as technical replicates. The model accounts for within-group variability to compare protein abundance. Simulations show LMM outperforms MS1- or MS2-only analyses, achieving higher significant ratios (SR) for true positives and lower SRs for false positives across various sample sizes. We validated this using urine-derived extracellular vesicles from pancreatic cancer patients and healthy controls (Orbitrap Eclipse/Spectronaut 19.1). The LMM identified more differentially abundant proteins (DAPs) than MS1-only but fewer than MS2, providing a balanced result that avoids t-test overestimation. This integration also improved pathway enrichment, offering a robust, interpretable framework for DIA-based proteomics.
Keywords
Data-Independent Acquisition (DIA)
Linear Mixed-Effects Model
mass spectrometry MS1/MS2
proteomics
Pancreatic cancer
Urinary extracellular vesicles (EVs)
In GSK's work developing oncology therapies, comprehensive characterization of the tumor microenvironment, particularly the interplay of target heterogeneity and expression, is informative of clinical potential, precise treatment strategies, novel biomarkers, and bispecific pairings. Traditional tumor characterization methods like immunohistochemistry are widely accessible due to their affordability and speed but are often restricted to analyzing a limited number of markers per sample. High-parameter spatial proteomic technologies address this challenge but are typically more costly and time-intensive.
This talk will present our approach to selecting regions of interest for high-plex spatial proteomics studies from higher-throughput, lower-plex studies. We used multivariate analysis to guide experimental design by implementing dimensionality reduction and clustering methods to select representative regions from a broad, lower-resolution imaging study. We demonstrate that exploratory studies can be leveraged for identification of diverse and characteristic samples to optimize learnings from spatial omics studies, reducing the cost and time needed for impact on drug discovery.
Keywords
spatial omics
experimental design
drug discovery
Single-cell multiome assays jointly profile gene expression (GEX) and chromatin accessibility (ATAC) in the same cell, enabling direct interrogation of chromatin-to-transcription relationships. Yet GEX and ATAC can partially decouple at the cell level, motivating open questions about how to characterize cross-modality concordance and leverage paired measurements for downstream analyses. Here we present POLARIS, a principled spectral framework with three main outputs: (i) cell-specific cross-modality concordance scores whose regimes correspond to agreement versus divergence between GEX and ATAC neighborhood structure, (ii) a joint cell embedding that refines cell type and state structure by combining separations from either modality, and (iii) a joint feature embedding that enables efficient genome-wide peak-gene association mapping. Together, POLARIS detects low-concordance cell populations enriched for modality-specific structure and recovers long-range peak-gene linkages that correlation-based screens often miss, while remaining computationally efficient and scalable to consortium-scale multiome datasets.
Keywords
single-cell multiome
multimodal integration
gene-peak joint embedding
cross modality concordance
spectral method
regulatory element to gene mapping
Speaker
Ziqi Fu, Harvard University
Co-Author(s)
Xihong Lin, Harvard T.H. Chan School of Public Health
Rong Ma, Harvard University