Contributed Poster Presentations: Biometrics Section

Tuesday, Aug 4: 2:00 PM - 3:50 PM
Contributed Posters 
Thomas M. Menino Convention & Exhibition Center 
Room: CC-Exhibit Hall A 

Main Sponsor

Biometrics Section

Presentations

01: A marginalized zero- and N-inflated binomial regression model for fractional outcomes

In health outcomes research, many clinical endpoints are fractional outcomes bounded in the [0,1] interval, such as medication adherence, measured as the proportion of prescribed doses taken. In many cases, fractional outcome distributions exhibit a high proportion of 0s and 1s, corresponding with non-engagement and consistent engagement, respectively. We developed a marginalized zero- and N-inflated binomial (MZNIB) regression model to capture a mixture distribution comprising structural 0s and 1s, and a binomial component for intermediate outcomes. Covariates in the MZNIB model are linked to the marginal mean via logistic regression, yielding straightforward, population-average interpretations. We developed score-based estimating equations derived from a working likelihood to estimate model parameters. Statistical inference can be made using a modified bootstrap approach, in which p-values are derived by inverting percentile bootstrap confidence intervals. Numerical studies were conducted to assess the feasibility and validity of the MZNIB model, which demonstrated good performance across varying data conditions. Real-data analyses also demonstrated satisfactory performance. 

Keywords

fractional outcome

marginal mean

floor and ceiling effect

marginalized zero- and N-inflated binomial regression model

mixture distribution 

Speaker

Zhengyang Zhou

Co-Author(s)

Minge Xie, Rutgers University
Eun-Young Mun, University of North Texas Health Science
Isaac Rhew, University of Washington
David Huh, University of Washington

02: A Trivariate Joint Modeling of Longitudinal, Recurrent and Terminal Events via Item Response Theory

We introduce a trivariate joint modeling framework that unifies longitudinal biomarkers, recurrent events, and terminal event risk through a latent disease‐severity process and a shared frailty component. The proposed model accommodates longitudinal outcome via an IRT formulation and captures complex relationships between disease progression and event processes. In this structure, a shared latent trait links a longitudinal outcome with both recurrent and terminal events, while a shared frailty term captures dependence between recurrent and terminal event risks. Together, these components allow the model to jointly characterize associations across longitudinal, recurrent, and terminal processes within a cohesive framework. Simulation studies show accurate parameter recovery and strong predictive performance under realistic study conditions, including sparse or irregular follow‐up. We demonstrate the utility of the approach using a cardiovascular disease dataset. 

Keywords

Joint Modeling

Bayesian Analysis

Item Response Theory 

Speaker

Dongrak Choi, Edwards Lifesciences

03: A Two-Stage Modeling Approach to Risk Prediction with High-Dimensional Two-Phase Data

When building risk prediction models using electronic health record (EHR) data, incorporating additional granular information from external sources may improve predictive accuracy. However, such external data is often available only for a subset of patients, resulting in monotone missingness. We formulate this problem within a two-phase design framework. In contrast to classical two-phase settings, the high dimensionality of EHR predictors and the complex availability mechanism of external data pose substantial challenges. To address these challenges, we propose a two-stage modeling method for building and evaluating risk prediction models for binary outcomes. In the first stage, we construct an initial prediction model using only EHR predictors under a sample-splitting procedure to summarize the high-dimensional predictors through a risk score. In the second stage, we fit a logistic regression model that recalibrates the initial model and incorporates the external variables as additional predictors. We develop a pseudo-score estimation approach that efficiently utilizes the EHR data and flexibly accounts for the differential availability of external data. We further propose a plug-in estimator for the area under the ROC curve. The proposed method is evaluated through theoretical characterization of its large sample properties and extensive simulation studies, and is further applied to develop a mortality risk prediction model for oncology patients using data from the University of Pennsylvania Health System EHRs enriched with additional patient survey information. 

Keywords

Electronic Health Records

High-dimensional risk prediction

Two-phase data

Two-stage Modeling 

Speaker

Chenyu Bi, Department of Biostatistics, Epidemiology and Informatics, University of Pennsylvania Perelman School of Medicine

Co-Author(s)

Jill Hasler
Changcheng Li
Ravi Parikh, Emory University, Winship Cancer Institute
Weidong Ma, Department of Biostatistics, Epidemiology and Informatics, University of Pennsylvania Perelman School of Medicine
Jinbo Chen, University of Pennsylvania

04: Accounting for Cellular Mixture in Spatially Aware Cell-Cell Interaction Analysis

Spatial transcriptomics (ST) enables genome-wide gene expression profiling while preserving tissue architecture, providing a powerful framework for studying cell–cell interactions (CCIs) in situ. However, many existing CCI methods inadequately account for mixed cellular composition within spatial measurements and often adapt single-cell frameworks without fully leveraging spatial information, and systematic benchmarking across platforms remains limited. We developed SpaCCI, a spatially aware CCI framework that models each spatial unit as a mixture of cell types by integrating ligand–receptor expression with location-specific cell-type abundance and neighborhood information to capture local and global interaction patterns. To contextualize SpaCCI, we benchmarked nine spatial CCI methods using realistic simulations and nine real datasets spanning Visium, Stereo-seq, and Xenium. Evaluations of prediction accuracy, spatial coherence, biological relevance, and computational efficiency revealed substantial performance variability across resolutions, tissues, and technologies, highlighting trade-offs and providing practical guidance for spatial CCI method selection and development. 

Keywords

Spatial Transcriptomics

Cell-cell Interaction

Ligand–receptor analysis 

Speaker

Li-Ting Ku, The University of Texas MD Anderson Cancer Center

Co-Author(s)

Vincent Bernard, Department of Gastrointestinal Radiation Oncology, The University of Texas MD Anderson Cancer Center
Jimin Min, Perlmutter Cancer Center, Department of Medicine, New York University Grossman School of Medicine,
Ying Yuan, University of Texas MD Anderson Cancer Center
Eugene J. Koay, Department of Gastrointestinal Radiation Oncology, The University of Texas MD Anderson Cancer Center
Anirban Maitra, Department of Pathology, New York University Grossman School of Medicine, NYU Langone Health
Liang Li, University of Texas MD Anderson Cancer Center
Ziyi Li, MD Anderson Cancer Center

05: Are machine learning methods for supervised classification truly better learners than classical LDA

The Increasing availability of large datasets with limited documentation or anonymized participant poses a serious threat to valid data analyses. Analyzing such data as an independent measurement without accounting for its longitudinal structure can bias results in some situations. To characterize this bias, we conducted simulations and analyzed repeated measures data as independent observations using Linear Discriminant Analysis (LDA) and machine learning (ML) algorithms: SVM, Random Forest, Neural Network, and K-Nearest Neighbor. We evaluated bias assuming independence for longitudinal datasets with compound symmetry & first-order autoregressive correlation structures, varying serial correlations (0–0.99), between-variable correlations (0.2–0.9), and variable variances. Results showed that when between-variable correlations were low, variances high, and serial correlation strong, ML models overestimated accuracy compared to LDA. LDA gave accurate results irrespective of the serial correlation levels among the observation. These results highlight the importance of understanding data generation and correctly modeling the data structure for valid inference, and reliable predictions. 

Keywords

Longitudinal Data

Machine Learning

Bias

Linear Discriminant Analysis

Compound Symmetry

AR(1) 

Speaker

Oluwafunmibi Fasanya, Louisiana State University Health Sciences Center

Co-Author

Kent Eskridge, University of Nebraska, Statistics Department

06: Bayesian Factored Regression for Efficient Analysis of Two-Phase Study Designs

In research settings where collecting expensive or resource-intensive covariates from all subjects is impractical, two-phase designs can substantially improve study efficiency compared to simple random sampling by strategically selecting participants based on outcome data and available covariates. However, accounting for the missing at random covariates that were not sampled is a key analytical challenge in two-phase studies. Failure to properly address this missingness can introduce bias in parameter estimates. We propose flexible Bayesian factored regression using MCMC to model the distribution of the missing covariates jointly with the main analysis regression model of interest. Key advantages include applicability to both continuous and discrete covariates, extensibility to various covariate distributions, and stability in model fitting with small samples. We compare the performance of our Bayesian method to the previously studied Ascertainment-Corrected Maximum Likelihood (ACML) approach through extensive simulation studies. We also illustrate our approach with a case study using data from the Lung Health Study and introduce an R package implementing these methods. 

Keywords

Two-phase designs

Bayesian factored regression

Missing data

Longitudinal data 

Speaker

Maximilian Rohde, Bristol Myers Squibb

Co-Author(s)

Jonathan Schildcrout, Vanderbilt University
Ran Tao, Vanderbilt University Medical Center

07: Bayesian Inference for Covariate-Adjusted Restricted Mean Survival Time Using Pseudo-Observations

Restricted mean survival time (RMST) is often used as an alternative summary measure of treatment effect to the hazard ratio, particularly when the proportional hazards assumption is violated. Although a frequentist regression model for estimating covariate-adjusted RMST based on pseudo-observations does not require modeling the survival function, this method often yields coverage rates below the nominal level in small-sample settings.
A Bayesian extension using the Bayesian generalized method of moments (GMM) generally provides coverage rates above the nominal level under the same settings. We proposed a method that modified Bayesian GMM from the perspective of general Bayesian inference, in which the uncertainty of the posterior distribution was adjusted by determining an appropriate learning rate. Numerical experiments assuming Weibull distributions revealed that the proposed method provided coverage rates closer to the nominal level than existing Bayesian methods in small-sample settings (n = 30, 50), while achieving coverage rates comparable to those of existing methods in large-sample settings (n = 200). 

Keywords

Survival analysis

Bayesian Inference

Restricted mean survival time

General Bayesian Inference 

Speaker

Ryoma Jibiki, Tokyo University of Science

Co-Author(s)

Tomohiro Ohigashi, Tokyo University of Science
Shunichiro Orihara, Tokyo Medical University
Takashi Sozu, Tokyo University of Science

09: Behavior of Confidence Intervals for Ridge Shrinkage Parameters in Beta Regression Models

Shrinkage estimators are widely used to address multicollinearity in regression models, yet their inferential properties are not well understood beyond classical linear settings. This paper investigates the behavior of confidence intervals of ridge shrinkage parameters in beta regression models. We conducted a large-scale Monte Carlo simulation study under varying sample sizes, predictor dimensions, and levels of correlation among covariates. The parameters are evaluated using mean squared error, confidence interval width, and coverage probability. The results demonstrate that shrinkage estimators can substantially reduce estimation error relative to maximum likelihood estimation under moderate to severe multicollinearity. However, estimators with stronger shrinkage often exhibit noticeable undercoverage, whereas moderate shrinkage rules tend to achieve a more favorable balance between efficiency and coverage. These findings highlight important trade-offs between bias reduction and inferential accuracy. The results are particularly relevant for methodological development and applied analyses. 

Keywords

Multicollinearity

Beta Regression

Ridge Regression 

Speaker

Sultana Mubarika Chowdhury, Florida International University

Co-Author(s)

Zoran Bursac, Florida International University
B.M. Kibria, Florida International University

10: Cross-Validation Bias in Presence-Only Spatial Models Applied to Topoclimatic Zoning

Random cross-validation can overestimate predictive performance in presence-only spatial models when spatial autocorrelation among clustered occurrences is ignored. We quantified this bias using 635 georeferenced occurrences of Carapa guianensis, a key species for Amazonian extractive economies. Random Forest models were fitted with WorldClim predictors under a factorial design combining two feature spaces, topoclimatic variables alone and topoclimatic variables plus geographic coordinates, and four validation schemes: random cross-validation, spatial block, environmental block, and kNNDM. Results showed that the coordinate-augmented model achieved an AUC of 0.96 under random cross-validation, but this value dropped to 0.53–0.62 under environmental and spatial blocking. In addition, a Local Join Count test on residuals detected statistically significant spatial clustering of classification errors (p < 0.05) not identified by global metrics. These results demonstrate that random cross-validation inflated model performance and that spatially structured validation provides more reliable support for topoclimatic zoning. 

Keywords

Spatial Cross-Validation

Spatial Autocorrelation

Spatial Leakage

Species Distrtibuition Models

Random Forest 

Speaker

Werlleson Nascimento, Universidade de São Paulo (ESALQ/USP)

11: Estimation of state occupation probabilities for the illness-death model via multiple imputation

In clinical research, multi-state models can provide a more granular description of disease progression than traditional outcomes such as overall survival or composite endpoints. State occupation probabilities for multi-state models can be derived from the Aalan-Johansen estimator, which also yields unbiased estimates of transition probabilities for data satisfying the Markov property. Recent work has employed multiple imputations of event times in right censored data to unbiasedly estimate survival and cumulative incidence functions, which are straightforwardly calculated as proportions after censored times have been imputed. In this talk, we extend these ideas to estimate state occupation probabilities and joint confidence regions for the three-state irreversible illness-death model, including when the Markov property does not hold. Through simulation, we empirically demonstrate that our imputation-based estimator unbiasedly estimates state occupation probabilities in both Markov and non-Markov settings, and corresponding pointwise confidence regions exhibit nominal coverage. 

Keywords

Multiple imputation

Survival analysis

Multi-state model

Illness-death model

Non-Markov model 

Speaker

Rachel Gonzalez, University of Michigan

Co-Author(s)

Walter Dempsey
Philip Boonstra, University of Michigan

12: From Mean to Quantiles: Post-hoc Quantile Calibration of Epigenetic Age Clocks

Quantiles of epigenetic age provide clinically meaningful information about heterogeneity in aging trajectories, enabling identification of individuals who age faster or slower than their peers beyond mean-based assessments. However, widely used epigenetic clocks primarily output mean predictions and do not directly provide quantile estimates or calibrated uncertainty. We propose a residual-based model switching framework that converts an existing mean epigenetic clock into a quantile clock without refitting the underlying methylation-to-age model. The method estimates global or conditional residual quantiles and yields calibrated quantile predictions of epigenetic age under mild assumptions. We provide theoretical justification, practical algorithms, and open-source software, and demonstrate the approach through simulations and applications to two real-world datasets across commonly used epigenetic clocks. 

Keywords

Biological age

Epigenetic age clocks

Quantile calibration

Conformal prediction

Prediction intervals 

Speaker

Menghan Yi

Co-Author(s)

Canyi Chen, University of Michigan
Peter Song, University of Michigan

13: Functional Fine-Mapping using Variational Inference

Fine-mapping identifies causal variants from genome-wide association study signals, but is challenging due to high correlation between genetic variants. This is fundamentally a variable selection in regression problem with highly correlated predictors. Functional annotations, such as transcription factor binding predictions, provide external information that can improve selection when properly integrated through prior distributions in a Bayesian framework. However, existing functional fine-mapping methods either assume known relationships between annotations and causality (requiring pre-specification) or employ two-stage procedures: first computing annotation-based priors separately, then performing fine-mapping conditional on these fixed priors. We introduce a Bayesian fine-mapping method that jointly learns annotation-causality relationships and performs variable selection in a unified single-stage framework. Inference is performed via variational inference where both the posterior distribution over causal configurations and the parameters governing the annotation informed prior are optimised simultaneously. We demonstrate our model's power through comprehensive simulations. 

Keywords

Fine-Mapping

Variational Inference

Deep Learning

Variable Selection 

Speaker

Lucas Siu

Co-Author

Marina Evangelou, Imperial College London

14: Future-Dependent Event Definitions in Time-to-Event Analysis: Bias and an Analytic Approach

Some time-to-event analyses use future-dependent event definitions, in which event occurrence can only be determined using future information. For example, diagnosis of certain conditions may require two positive tests, with event times retrospectively assigned to the first test. Individuals with a single positive by the end of follow-up are typically censored, without accounting for the possibility of future confirmation, which leads to bias in Kaplan-Meier estimation. We propose a clone-based weighting approach to address this uncertainty. For individuals censored after a single event, the method generates clones representing alternative future event status and weights them by the probability of future event occurrence. Simulation studies assessed bias from future-dependent outcomes and evaluated the proposed approach. The results show that such outcomes can induce substantial bias in Kaplan–Meier estimation, while the proposed clone-weighting approach mitigates bias. Overall, these findings highlight the importance of avoiding time-to-event outcomes defined using future information and suggest clone-based weighting as an analytic approach when such definitions are unavoidable. 

Keywords

Survival Analysis

Time-to-event data 

Speaker

Hongseok Kim, CSL Behring

15: Imputing Multiple Missing Mixed-Type Covariates for Accelerated Failure Time Models

Missing covariates are common in biomedical studies with survival outcomes. Multiple imputation by chained equations (MICE) is widely used due to its broad software support and modeling flexibility, but implementations often rely on convenient conditional models which may not be compatible with the joint distribution implied by the analysis model. For accelerated failure time (AFT) models, methods address this issue by constructing MICE conditionals from joint models and fit separate imputation models by failure status. However, existing methods are limited to at most two missing covariates and require multilevel categorical covariates to be impute through separate dummy variables, which may yield invalid combinations and compromise inference in multivariable settings. To address these limitations, we extend the existing AFT imputation strategies to allow multiple mixed-type covariates, including continuous and categorical variables of binary, ordinal, or nominal types from general location and logistic normal joint models. Numerical studies demonstrate that the proposed methods reduce bias, improve coverage, and achieve comparable efficiency to conventional imputation strategies. 

Keywords

Censored survival data

Accelerated failure time model

Missing covariates

Multiple imputation by chained equations

Mixed type covariates 

Speaker

Siyao Li

Co-Author(s)

Lihong Qi, UC Davis
Yulei He, AbbVie
John Robbins, Department of Internal Medicine, School of Medicine, University of California, Davis, USA
Shuai Chen, UC Davis-Public Health Sciences

16: Integrated Likelihood CI for Binomial Proportion using Double Sampling under Misclassification

We consider the problem of interval estimating a binomial proportion parameter when data are subject to under-reporting using a double-sampling design. This double-sampling framework enables identifiability of all model parameters by incorporating additional information that increases the dimension of the sufficient statistics relative to the number of parameters. We focus on constructing two new confidence intervals (CIs) for the target proportion. We then contrast the performance of four current CIs with nuisance parameter and two proposed integrated likelihood (IL) CIs. Through extensive Monte Carlo simulations and two real-data examples, we evaluate the coverage probability and interval width of each method for many combinations of parameter settings and sample sizes. We demonstrate that an IL approach offers better nominal coverage and more stable CIs than five competing CIs. 

Keywords

Binary Data

False Positive Observations

Pseudo-Likelihood Methods 

Speaker

Terry Tsai

Co-Author(s)

Dean Young
Rakheon Kim, Baylor University

17: MEM-Seq: A scalable adjustment method for correcting spatial autocorrelation in spatial multiscale analysis of microbiome sequencing data

Recent advances in spatial profiling technologies have enabled the collection of microbiome sequencing data with explicit spatial information, creating new opportunities—and challenges—for microbiome analysis. In this study, we introduce and analyze a novel data type: spatially resolved microbiome sequencing data obtained from a mouse intestinal tissue section containing a tumor. Samples were collected from multiple spatial locations across tumor and adjacent normal regions, with the goal of characterizing microbial community differences and identifying factors driving these differences. A key methodological challenge in spatial microbiome analysis is spatial autocorrelation, which violates the independence assumptions underlying many standard statistical methods. We developed a novel and principled approach for addressing the spatial autocorrelation by borrowing and improving ideas from classic spatial statistics toolkits. Our approach allows for scalable and effective analysis on the microbiome community level differences between tissue compartments (such as normal vs tumor). Applying our method, we identify significant microbial community differences between tumor and normal tissue regions after accounting for spatial autocorrelation. Notably, we find that sequencing depth emerges as a key factor driving community variation, challenging the common assumption that sequencing depth acts solely as a technical nuisance variable. These findings highlight the importance of jointly modeling spatial structure and experimental factors in spatial microbiome studies and provide new insights into tumor–microbiome interactions, laying the groundwork for future methodological and biological investigations in spatial microbiome analysis. 

Speaker

Daxuan Deng

18: Microbiome data analysis using a zero-inflated logistic normal multinomial model with covariates

Microbial abundance data, called microbiome data, are characterized by zero-inflated count structures. Zeng et al. (2022) proposed the zero-inflated logistic normal multinomial (ZILNM) model. A key characteristic of this model is the introduction of latent factors, representing unobserved variables that influence microbial compositions.

In this study, we extend the ZILNM model by incorporating observed subject-level covariates in addition to latent factors. While both latent factors and covariates affect model parameters, covariates are directly observable and can represent clinically meaningful information. By including covariates, the proposed model is expected to provide a more interpretable and practically relevant framework for medical microbiome data analysis.

Simulation studies were conducted under settings where latent factors were present and subjects were divided into control and treatment groups. Variational Bayesian inference was applied for parameter estimation, and we examined the behavior and estimation performance of regression coefficients associated with covariates, with particular emphasis on group comparison and interpretability. 

Keywords

zero-inflated models

count data

latent factors 

Speaker

Yuki Ando

20: Modified Causal Forest for Heterogeneous Treatment Effects with Structurally Constraining Covariates

To improve generalizability of clinical trial results, studies may allow participation from individuals ineligible for one study treatment, as in the Biomarkers for Evaluating Spine Treatment (BEST) Trial. In this setting, challenges arise when estimating heterogenous treatment effects (HTEs) for a candidate biomarker that also influences treatment eligibility. This biomarker may simultaneously act as a prognostic factor, treatment effect modifier, and impact treatment assignment. Standard HTE approaches fail to account for these joint mechanisms, leading to biased estimates for the target population. We propose a framework that integrates inverse probability weighting with causal forests to address threats to generalizability arising from structural randomization constraints. This approach performs a dual adjustment for differences across eligibility subgroups and restricted treatment randomization, ensuring that estimated conditional average treatment effects (CATEs) are representative of the target population. We discuss the estimator's theoretical properties, show improved performance through simulations, and compare it to the standard causal forest approach using data from the BEST Trial. 

Keywords

Precision Health

Heterogeneous Treatment Effects

Causal Forest

Clinical Trial

Generalizability

Causal inference 

Speaker

Annika Cleven, University of North Carolina at Chapel Hill

Co-Author(s)

Bryce Rowland, University of North Carolina at Chapel Hill
Michael Kosorok, University of North Carolina at Chapel Hill
Kevin Anstrom, UNC-Chapel Hill
Matt Mauck, University of North Carolina - Chapel Hill

21: ReFIT: Federated Transfer Learning for Sequential Prediction and Uncertainty Quantification Using Streaming EHR Data

Modern biomedical data are increasingly collected across multiple institutions and time periods, creating opportunities for improved statistical inference through knowledge transfer, but also posing challenges for privacy, scalability, and distributional heterogeneity. We propose a Renewable Federated Incremental Transfer framework, termed ReFIT, for sequentially integrating information from streaming source datasets to improve model estimation and prediction in a target population. ReFIT builds upon a density ratio model to account for covariate shift between the source and target populations and employs a renewable updating strategy that allows model parameters to be incrementally refined as new source data become available, using only summary-level information from prior sources. This framework ensures privacy preservation and computational efficiency while adapting to evolving data environments. Beyond improving predictive performance, ReFIT also quantifies predictive uncertainty within a conformal prediction framework, yielding valid prediction intervals that adapt as new information accumulates. Extensive simulation studies demonstrate that ReFIT achieves higher predictive accuracy and better uncertainty quantification than models trained on target or source data alone. The method remains robust even under nonlinear model misspecification and varying degrees of source–target shift. Moreover, as ReFIT incrementally integrates additional source data, the conformal prediction intervals become progressively narrower without sacrificing coverage, evidencing improved statistical efficiency with growing information. In an electronic health record application for breast cancer prediction, ReFIT substantially improves prediction for the Hispanic population by sequentially leveraging information from non-Hispanic White patients collected over multiple time periods. These results highlight the potential of ReFIT as a general and practical framework for privacy-preserving, adaptive, and scalable learning from distributed and periodically updated biomedical data. 

Speaker

YUYING LU

Co-Author(s)

Lan Luo, Department of Biostatistics and Epidemiology, Rutgers School of Public Health
Tian Gu, Columbia University

22: Stochastic ordering of extreme first passage times for simple neural models

S.E. Lawley (2020, J. Math. Biol. 80:2301-2325) and subsequent papers examined the distribution of extreme first passage times to a constant threshold for a variety of diffusion processes. He starts with an ensemble of population of identiacal diffusion processes and considers the fastest one or minimum among the population to reach the threshold. He called these "extreme first passage times". Here we use previous results on the almost sure stochastic ordering of first passage times for one dimensional diffusion processes (L. Sacerdote, C. E. Smith, 2004, Methodology and Computing in Applied Probability 6:323-341) to examine stochastic orderings in extreme first passage times for several commonly used neural integrate and fire models. 

Keywords

Gumbel distribution

diffusion processes

Neural models

first passage times 

Speaker

Charles Smith, North Carolina State Univ.

23: Uncovering Hidden Spatial Structure Across 3D Biofilm Images

Microscopic images of microbial communities often reveal intricate spatial organization, where different species form structured spatial "neighborhoods" that are likely to reflect cooperation, competition, and shared function. A striking example is the biofilm community on the human tongue, where microbes arrange themselves in highly organized patterns at micrometer scales. Understanding these spatial relationships is key to understanding biological function—but statistically modeling them is challenging, especially when each sample gives rise to many biofilm images or slices.

Multivariate log-Gaussian Cox processes are flexible models for the analysis of multivariate point patterns. However, they have so far been focused on single realizations only (i.e., single images), ignoring similarity and dissimilarity across images. In this work, we develop a hierarchical Bayesian multivariate log-Gaussian Cox process framework that allows us to investigate both the common spatial interaction patterns shared across repeated images and how individual image slices differ from one another. By modeling all images jointly and borrowing information across slices, the approach improves the ability to detect meaningful interaction patterns while accounting for image-to-image variability.

Using simulation studies, we demonstrate accurate recovery of multitype spatial relationships and effective sharing of information across images. We also apply the framework to 3D Z-stack images of human tongue dorsum biofilms. By moving beyond post hoc slice-by-slice comparisons, the proposed approach enables integrated inference on shared spatial structure and image-specific variation in multivariate, multilevel biofilm imaging data. 

Speaker

Shuwan Wang

24: Using Linear Mixed Models to Assess Lactate Response to Exercise in Alzheimer's Disease vs. Controls

We studied exercise on brain glucose metabolism and lactate availability in individuals with Alzheimer's disease (AD) and controls. Whole-blood lactate was measured at six points during rest and exercise conditions. We modeled lactate response at rest and exercise across two intensity levels (moderate, higher). Linear mixed models were used to assess the relationship, accounting for repeated measures within subjects using random intercepts. We explored the covariance structure of the model and accommodated heterogeneous variances between rest and exercise. Estimated variances varied greatly, from 0.02 at rest to 1.20 after 20 minutes of exercise. Assuming no differences in baseline fitness level, we found a significant difference in lactate between AD and controls in the higher intensity exercise group. When accounting for baseline fitness, this difference is not significant. Our model suggests that those with AD have lower lactate availability under higher intensity exercise than controls. It is unclear if this is due to characteristics of AD or differences in baseline fitness. Future studies should more carefully control for baseline fitness to further explore this relationship. 

Keywords

Linear Mixed Models

Repeated Measures

Heterogeneous Variances

Alzheimer's disease 

Speaker

Lauren Yoksh

Co-Author(s)

Jonathan Mahnken, University of Kansas Medical Center
Jill Morris, University of Kansas Medical Center
Zachary Green, University of Kansas Medical Center

25: Regression Analysis of Misclassified Current Status Data with Potentially Unknown Test Accuracy

Current status data are frequently encountered in many real life cross-sectional epidemiological and demographic studies, where each subject is examined only once, and the failure time of interest is never exactly observed but known to be either smaller or larger than the examination time for each subject by evaluating the failure status. Consequently, current status data are a mixture of left-censored and right-censored observations for the failure times of all subjects with or without covariates. In some real-life studies, the test or diagnosis that is used to determine the failure status may be error-prone, and this leads to misclassified failure status for some or all subjects. The resulting data are referred to as misclassified current status data in the literature. We study regression analysis of misclassified current status data and propose a novel estimation approach under the proportional odds model. Specifically, monotone splines are adopted to approximate the baseline odds function, and an efficient expectation-maximization algorithm is developed based on a data augmentation involving exponential and Poisson latent variables.
We also develop an extension of the proposed method to account for unknown test accuracy. Simulation studies demonstrate excellent estimation performance, and the proposed method is illustrated through an application to uterine fibroid data. 

Speaker

Zhixin Chen, Bristol Meyers Squibb

56: Leveraging Bayesian Hierarchical Model to Combine Domain Knowledge and Big Data to Accelerate Crop Improvement

Sustained genetic gain in plant breeding requires accurate prediction of varietal performance across diverse environmental and management conditions. Contemporary predictive approaches typically fall into two classes: process-based crop models that encode mechanistic knowledge of plant development, and flexible data-driven models that emphasize predictive accuracy but offer limited interpretability. We propose a hybrid modeling framework that integrates these paradigms within a Bayesian hierarchical model. Outputs from a mechanistic crop growth model are incorporated as structured components alongside genomic covariates, allowing biological knowledge to inform inference while retaining statistical flexibility. The hierarchical structure enables partial pooling across environments and explicit uncertainty quantification.
Model behavior and predictive performance are evaluated using both simulated data and large-scale empirical breeding trial datasets. Results show that the proposed hybrid approach improves prediction and interpretability relative to purely data-driven alternatives, illustrating the value of combining process-based and statistical modeling in high-dimensional agricultural prediction problems.
 

Speaker

Tom Tang, Corteva Agriscience

Withdrawn: 08 Bayesian Recursive Copula Survival Models with Functional Covariates

Causal inference for time-to-event outcomes faces challenges from endogenous treatments and high-frequency functional covariates. While recursive copula models address endogeneity via latent dependence, existing methods struggle to jointly model binary exposures and survival endpoints with high-dimensional functional data. We propose a Deep Probabilistic Recursive Copula Model integrating variational autoencoders (VAEs) with recursive copula-based survival modeling. Our framework learns nonlinear representations of functional profiles via VAEs to parameterize the copula structure, enabling scalable modeling of complex exposure–outcome dependence. By employing variational inference, we approximate the posterior via evidence lower bound maximization, avoiding computationally intensive MCMC. Simulations show our approach yields lower bias than two-stage estimators and significant computational gains over MCMC-based Bayesian models. We apply our approach to REGARDS and NHANES 2012-2024 to estimate the causal effect of chronic disease status on survival adjusting for physical activity behavior.  

Keywords

Recursive Copula

Variational Inference

Functional Data Analysis

Survival Analysis

Endogeneity

Deep Learning 

Co-Author(s)

Roger Zoh, Indiana University
Carmen Tekwe, Indiana University

Withdrawn: 19 Missing data strategy in bioavailability studies at Danone Research & Innovation

Missing data are common in bioavailability studies with dense sampling. In the current approach, at Danone Research & Innovation, if a value is missing near the expected peak, the visit is marked unevaluable and a replacement participant recruited; otherwise simple interpolation may be used. This can reduce power, introduce bias in Cmax (peak concentration), Tmax (time to peak), and iAUC (incremental Area Under the Curve), and increase cost and burden. We aim to retain as many analysable visits as possible without compromising inference. We are developing a Bayesian multiple‑imputation workflow that models the time course jointly, borrowing strength from neighbouring time points, visit indicators, design covariates, and (optionally) historical healthy‑volunteer datasets. Priors allow nonlinear trends, and imputation uncertainty propagates to endpoints. We present simulations reflecting internal trials showing impacts on Cmax, Tmax, and iAUC, with diagnostics, MAR and plausible MNAR sensitivity and early rules on when imputation supports analysis without replacing participants. This pragmatic approach seeks to protect inference, reduce waste, and remain aligned with good practice. 

Keywords

Missing data

Bayesian

Imputation

Clinical trials

Bioavailability 

Co-Author(s)

Floor van Oudenhoven, Danone Research & Innovation
Sophie Swinkels, Danone Research & Innovation