Thursday, Aug 6: 8:30 AM - 10:20 AM
6004
Contributed Papers
Thomas M. Menino Convention & Exhibition Center
Room: CC-203
Main Sponsor
Biometrics Section
Presentations
Huntington disease (HD) is a genetically inherited neurodegenerative disease with progressively worsening symptoms. Accurately modeling time to HD diagnosis is essential for clinical trial design and treatment planning. Langbehn's model, the CAG-Age Product model, the Prognostic Index Normed model, and the Multivariate Risk Score model have all been proposed for this task. Because they differ in methodology, assumptions, and accuracy, these models may yield conflicting predictions. Previous model comparisons are limited and those that exist could be misleading due to (i) testing the models on the same data used to train them and (ii) failing to account for high rates of right censoring (80%+) in performance metrics. We discuss the theoretical foundations of these models, offering comparisons about their practical feasibility. Further, we externally validate their risk stratification abilities using data from the ENROLL-HD study and performance metrics incorporating inverse probability of censoring weights and Kaplan-Meier adjustments. We also show these models can be used to estimate sample sizes for an HD clinical trial, emphasizing that estimates from previous work would lead to underpowered trials.
Keywords
Concordance index
Censored covariates
Cox model
Inverse probability of censoring weighting
Neurodegenerative disease
ROC curves
Joint modeling of longitudinal data and survival data has been extended to accommodate multilevel data structures. In dental studies, data often exhibit a multilevel hierarchy: each patient has multiple teeth, and one or more biomarkers are measured repeatedly over time for each tooth. In addition to biomarker measurements, the time to tooth loss may vary differently between patients as some patients are more susceptible to tooth loss, conditional on other risk factors. In this paper, we account for intra-patient and intra-tooth correlations in the longitudinal measurement of a continuous biomarkers, probing pocket depth (PPD), and a binary biomarker, mobility. We also account for the correlation in time to tooth loss between teeth within the same patient. We jointly model the two biomarker measurements and the risk of tooth loss using a Bayesian estimation approach. We develop a cluster-informed dynamic prediction framework for the survival outcome, in which the prediction for a given unit is informed not only by its own observed history, but also by the observed longitudinal trajectories and event outcomes of other units within the same cluster. We evaluate the predictive performance in terms of discrimination and calibration, accounting for censoring and multilevel data structure. Our simulation study shows that the proposed joint model produced more desirable estimates and better predictive performance compared to the standard bivariate joint model that ignores the multilevel data structure. We applied our model to electronic periodontal data obtained from the Canadian Armed Forces (CAF).
Keywords
Multilevel data
Joint model
Dynamic Prediction
Discrimination
Calibration
Periodontitis
Measures of explained variation, such as prognostic R², are widely used to evaluate survival models and to inform key design decisions, including sample size planning and external validation. In time-to-event settings, explained variation is not solely determined by the prognostic signal, but also by the amount of event information observed. We show that prognostic R² measures are inherently design-dependent and require explicit conditioning on censoring and recruitment mechanisms. Using controlled simulations with fixed hazard ratios, we investigate the behavior of common R² statistics under varying administrative censoring and follow-up. Likelihood-based measures show substantial downward bias as censoring increases, even when prognostic separation is unchanged, whereas separation-based measures remain stable. These results indicate that differences in explained variation may primarily reflect study design rather than model performance. We discuss implications for comparative modeling, external validation, and planning of prognostic studies, emphasizing that explained variation depends on both prognostic signal and information accrual.
Keywords
Survival analysis
Study design
Censoring
Prognostic R²
Prognostic models
Risk prediction
Speaker
Ester Rosa, University of Padua, Italy
Co-Author(s)
Stefania Lando, University of Padova
Gloria Brigiari, Unit of Biostatistics, Epidemiology and Public Health Department of Cardiac, Thoracic, Vascular Sciences, and Public Health University of Padova
Dario Gregori, University of Padova
We develop a semiparametric prediction method for outcomes with right-censored covariates. Fixing a nominal coverage level and an interval center, we estimate the interval length via a semiparametric approach using an efficient influence function. The estimator is semiparametrically efficient, yielding highly accurate prediction intervals, and is doubly robust to misspecification of nuisance models, allowing flexible choices for the nuisance models. Moreover, we establish asymptotic validity for the resulting prediction coverage. Simulations and a Huntington disease application show that our method produces prediction intervals with more stable interval lengths and prediction coverage closer to the nominal level, compared to distribution-free baseline methods such as conformal prediction.
Keywords
semiparametric modeling
double robustness
semiparametric efficiency
censored covariates
conformal prediction
Huntington disease
Longitudinal time-to-event studies often index covariates to an imprecise subject-specific origin (e.g., estimated disease onset or conception), creating an unknown additive shift in individual time scales. Standard landmark and joint models that use the observed clock as correct can yield biased dynamic risk estimates. We introduce SILK (Shift-Invariant Landmark Kernels), a methodology for dynamic prediction under subject-specific time-shift error. SILK reformulates landmark analysis on shift-invariant primitives-residual time-to-event and visit increments-so that estimation and prediction do not require specifying a parametric measurement-error model. At each landmark, SILK combines (i) a survival model for the residual-time outcome with (ii) nonparametric learning of the conditional distribution of longitudinal biomarkers given history via RKHS conditional mean embeddings that capture evolving biology. We establish identifiability of the latent time shift when biomarkers encode stage information and provide finite sample guarantees. Simulations show gains under realistic misalignment. Application to pregnancy ultrasound biomarkers shows improved dynamic PTB risk prediction.
Keywords
Survival Analysis
Measurement Error
Reproducing Kernel Hilbert Space
Non-parametric
Longitudinal data
Radiation therapy induces profound disruptions in lymphocyte homeostasis, yet prognostic assessment in clinical practice largely relies on simple summaries such as absolute lymphocyte count (ALC) nadirs, which overlook rich temporal information encoded in the full ALC trajectory during and after chemoradiation. We propose a stochastic modeling framework to quantitatively characterize patient-specific radiation-induced lymphocyte kinetics. Built on state-dependent multi-type branching processes and diffusion approximations, this two-phase mechanistic model captures treatment-driven cytotoxicity and post-treatment regenerative dynamics through distinct self-renewal, death, and egress rates. Application to longitudinal ALC data from esophageal cancer patients receiving chemoradiation shows the estimated individual-level kinetic parameters from ALC trajectories provide prognostic information beyond nadir-based metrics. These results highlight the proposed framework as a mechanistically interpretable modeling of radiation-induced lymphocyte kinetics with potential to inform risk stratification and treatment planning in radiation oncology.
Keywords
Stochastic modeling
Longitudinal data
Oncology
Radiation therapy
Survival analysis
Immune kinetics
Speaker
Radhe Mohan, Department of Radiation Physics, Division of Radiation Oncology
Co-Author(s)
Yiqing Chen
Ren-Yi Wang
Steven Lin, Division of Radiation Oncology, The University of Texas MD Anderson Cancer Center, Houston, TX
Evaluating and validating the performance of prediction models is a fundamental task in statistics, machine learning, and their diverse applications. However, developing robust performance metrics for competing risks time-to-event data poses unique challenges. We first highlight how certain conventional predictive performance metrics, such as the C-index, Brier score, and time-dependent AUC, can yield undesirable results when comparing predictive performance between different prediction models. To address this research gap, we introduce a novel time-dependent pseudo R^2 measure to evaluate the predictive performance of a predictive cumulative incidence function over a restricted time domain under right-censored competing risks time-to-event data. Specifically, we first propose a population-level time-dependent pseudo R^2 measures for the competing risk event of interest and then define their corresponding sample versions based on right-censored competing risks time-to-event data. We investigate the asymptotic properties of the proposed measure and demonstrate its advantages over conventional metrics through comprehensive simulation studies and real data applications.
Keywords
Competing risks
Prediction performance
Survival models
Explained variance
C-index
Brier Score