Contributed Poster Presentations: ENAR

Tuesday, Aug 4: 2:00 PM - 3:50 PM
Contributed Posters 
Thomas M. Menino Convention & Exhibition Center 
Room: CC-Exhibit Hall A 

Main Sponsor

ENAR

Presentations

31: DISCO: Diagnosis of Separation and Correction of Odds-ratio Inflation in Logistic Regression

Logistic regression models binary outcomes in biomedical studies to obtain interpretable odds ratios. Separation occurs when predictors perfectly or nearly perfectly classify cases and controls, causing maximum likelihood estimates to diverge. We develop DISCO (DIagnosis of Separation and Correction of Odds-ratio inflation), a detection-and-estimation framework. DISCO provides pre-hoc diagnosis to detect separation, identify problematic predictors, and quantify severity. A key theoretical insight is that under separation, the coefficient vector's direction remains identifiable, so pairwise ratios of nonzero coefficients are well-defined despite individual magnitudes diverging. We propose a Bayesian estimator with multivariate exponential power prior that penalizes coefficients. We construct a simulation framework generating separation cases with ground truth for fair comparison. Simulations show DISCO detects more problematic data than existing methods, and our estimator controls bias better than alternatives. In HIV-risk data, DISCO outperforms competitors by identifying problematic variables and producing finite, sign-preserving odds ratios. 

Keywords

Logistic regression

Separation Problem

Bayesian methods

Multivariate exponential-power prior 

Speaker

Chenyu Liu, Case Western Reserve University

Co-Author(s)

Zihan Zhu, Yale school of public health, Department of biostatistics
Xi Qiao, Huntsman Cancer Institute at the University of Utah
Liangliang Zhang, Case Western Reserve University

32: Evaluation of Rater Performance for Ordinal Diagnostic Tests Using Latent Class Models

Many ordinal discrete diagnostic tests, such as breast cancer tumor grade, are prognostic factors but are often subjective and lack reproducibility. With multiple independent ratings, it is important to evaluate and compare performance among raters.. Here first we consider testing whether one of the raters can act as the gold standard, with the alternative that the gold standard is latent and the observed ratings are modeled by a latent class model (Kruskal, 1977; Dawid & Skene, 1979). When indeed none of the raters can act as the gold standard, we consider both the maximum likelihood method and the Bayesian method for estimating and comparing the concordance between each rater and the latent gold standard. Both simulation studies and application to tumor grade and molecular subtypes of breast cancer are used to illustrate these methods. 

Keywords

Discrete diagnostic tests

Latent class models 

Speaker

Xinyi Zhang, The University oF pittsburgh

Co-Author

Gong Tang, University of Pittsburgh

33: Machine Learning Evaluation using Semiparametric Correlation in Brain-Psychopathology Associations

Machine learning (ML) models are used in neuroimaging studies to predict biological and psychometric phenotypes, such as age and psychopathology factor scores. Neuroscientists use Pearson's correlation between the predicted and actual feature to quantify model accuracy for these ML models; however, Pearson's correlation is not accurate when using ML models due to their slow. To address this, we use a model-agnostic semiparametric "one-step" (OS) modification of Pearson's correlation to model associations between functional and structural neuroimaging data with age and psychopathology in the Reproducible Brain Charts (RBC) dataset. We use random forest and ridge regression machine learning models and are able to provide model accuracy scores with confidence intervals that are less biased and more replicable. Our method allows valid model comparisons not possible with Pearson's correlation and shows that functional and structural imaging are highly predictive of age, but they do not improve prediction accuracy for psychopathology. 

Keywords

Machine Learning

Psychopathology

Semiparametric Correlation

Prediction 

Speaker

Ishaan Gadiyar, Vanderbilt University Medical Center

Co-Author(s)

Megan Jones
Simon Vandekar, Vanderbilt University Medical Center

34: Recovering Missing Correlations in Multi-Site Metabolomic Studies

Many NIH-sponsored laboratories carry out nontargeted metabolite studies, focusing only on a subset of the entire human metabolome and reporting summary data for the subset of metabolites. To form a complete picture of the entire metabolome or a substantial part of it, one must put the summary analysis on different subsets, obtained from the different laboratories, together in a consistent and efficient manner. We propose a methodology for integrating findings at various laboratories in a cooperative research partnership, in a scenario where only estimated Spearman rank correlation matrices, not raw data, can be shared between partner organizations. We assume the laboratories will study different, but intersecting sets of metabolites and consider how to impute missing values for the metabolite data not collected at a given lab. 

Keywords

meta-analysis

metabolomics

distributed research

correlation matrix 

Speaker

Ryan Lafferty, University of Maryland, Baltimore County

Co-Author(s)

Anindya Roy, University of Maryland-Baltimore County
Paul Albert, National Cancer Institute

35: SVD-PICAR Bayesian Kernel Machine Regression for Spatio-Temporal Count Data

Bayesian kernel machine regression (BKMR) has become a widely used tool for studying the health effects of complex exposure mixtures, with extensions to over dispersed count outcomes now available. However, existing BKMR models for count data handle spatial and temporal structure separately rather than as a unified random field, leaving the joint space-time dependency in the data unmodeled. No existing method therefore combines nonlinear mixture modeling, a neighborhood-based spatio-temporal random effect, and an over dispersed count likelihood in a single computationally workable framework.

We address this gap by proposing SVD-PICAR NB-BKMR, which extends the PICAR framework from the spatial to the spatio-temporal setting through a Kronecker product of spatial and temporal bases. However, this spatiotemporal basis contains redundant components that expand the dimension of the latent random field, produce highly correlated posterior draws, and make the full-rank ICAR field ill-conditioned, weakening exposure-effect estimates and slowing MCMC sampling. We resolve this by incorporating a truncated singular value decomposition (SVD) into the spatio-temporal PICAR basis, retaining only the dominant orthogonal directions, reducing the latent field to a manageable dimension, and separating background spatio-temporal variation from the nonlinear exposure-response pattern.
We compare three competing models: a fixed effect model (M1), a full-rank mixed effect model (M2), and the proposed SVD-PICAR mixed effect model (M3). Models are evaluated through simulation and an application to county level thyroid cancer and pesticide exposure data from Nebraska (1992–2014). In simulation, M3 achieves the best predictive fit by the Watanabe-Akaike Information Criterion (WAIC), the highest MCMC efficiency by effective sample size per second (ESS/sec), and the most stable exposure identification. In the Nebraska application, Atrazine and Alachlor emerge as the leading contributors to thyroid cancer risk, with nonlinear exposure-response relationships that remain robust after spatio-temporal correction. The SVD-PICAR framework offers a principled and practically workable approach for mixture analysis in spatio-temporal count data settings.
 

Keywords

Spatio-temporal count data


Bayesian kernel machine regression

projection based intrinsic conditional autoregression

singular value decomposition

basis representation

negative binomial 

Speaker

Kyei Afari, University of Nebraska Medical Center