AI, Machine Learning & Digital Tools in Clinical Development

Santosh Sutradhar Chair
Merck & Co., Inc.
 
Monday, Aug 3: 10:30 AM - 12:20 PM
6030 
Contributed Papers 
Thomas M. Menino Convention & Exhibition Center 
Room: CC-206B 

Main Sponsor

Biopharmaceutical Section

Presentations

A transformer-based model to predict disease course of newly diagnosed and relapsed multiple myeloma

Multiple myeloma management requires a balance between maximizing survival, minimizing adverse events to therapy, and monitoring disease progression. We developed a transformer-based machine learning model that jointly (1) predicts progression-free survival (PFS), overall survival (OS), and adverse events (AE), (2) forecasts key disease biomarkers, and (3) assesses the effect of different treatment strategies, e.g., ixazomib, lenalidomide, dexamethasone (IRd) vs lenalidomide, dexamethasone (Rd). Using TOURMALINE trial data, we trained and internally validated our model on newly diagnosed myeloma patients (N = 703) and externally validated it on relapsed and refractory myeloma patients (N = 720). Our model achieved superior performance to a risk model based on the multiple myeloma international staging system (ISS) and comparable performance to survival models trained separately on each task, but unable to forecast biomarkers. Our approach outperformed state-of-the-art deep learning models, tailored towards forecasting, on predicting key disease biomarkers. 

Keywords

Artificial Intelligence

Transformer Model

Multiple Myeloma

Clinical Trial 

Speaker

Cong Li

Co-Author(s)

Zeshan Hussain, CSAIL, MIT
Edward Brouwer, CSAIL, MIT
Rebbeca Boiarsky, CSAIL, MIT
Sama Setty, CSAIL, MIT
Neeraj Gupta, Takeda Development Center Americas
Guohui Liu, Takeda Pharmaceuticals International Co.
David Sontag, CSAIL, MIT

An Ensemble Transfer Learning Framework for Robust Polygenic Prediction of Drug Response

Accurate prediction of drug response in pharmacogenomics (PGx) is hindered by traditional methods, such as disease-specific polygenic risk scores (PRS-Dis), which often fail to capture the genetic basis of drug efficacy. Direct PGx PRS approaches could improve predictions, but their application is limited by the lack of large, relevant PGx datasets. To address these challenges, we present PRS-PGx-ETL, a novel Ensemble Transfer Learning (ETL) framework that combines transfer learning (TL) and ensemble learning (EL) to improve drug response prediction. TL leverages genetic data from large disease or PGx-T (treatment-only) base cohorts and transfers knowledge to a target PGx cohort. EL integrates multiple PRSs, generated from different base cohorts and various PRS methods, into a single, optimally weighted score. This integrative approach expands the genetic information used and automates parameter tuning. In simulations and application to the IMPROVE-IT PGx GWAS dataset, PRS-PGx-ETL significantly improves drug response prediction accuracy and patient stratification over existing methods. Our framework offers a robust PRS tool for advancing precision medicine. 

Keywords

GWAS

Pharmacogenomics

Polygenic risk score

Transfer learning

Ensemble learning 

Speaker

Youshu Cheng

Co-Author

Judong Shen, Merck & Co., Inc.

Bayesian Machine Learning Counterfactual Treatment Selection for Survival Outcome

Background: Atrial fibrillation (AF) carries high morbidity, and response to catheter ablation varies widely. Existing personalized treatment approaches focus on simple outcomes, while survival and time‑dependent treatment effects are rarely considered.
Method: We propose a Bayesian causal model to identify patient profiles showing differential treatment benefit for time‑to‑event outcomes. Using DECAAF II data, we modeled time to AF recurrence with a survival random forest using baseline covariates and treatment. Synthetic patient profiles were generated by bootstrap resampling, and counterfactual survival predictions were obtained for both treatments. Benefit at 270 days was defined as the difference in predicted survival risk. These differences were used to build an interpretable treatment‑effect tree. Real trial participants were then projected onto subgroups and annotated with observed 270‑day cumulative incidence and censoring rates.
Evaluation: Evaluation used a model‑based mortality index, defined as the difference in cumulative hazard at 270 days. Real patients showed subgroup‑specific patterns consistent with predictions, identifying groups favoring each treatment. 

Keywords

Bayesian prediction

Individualized treatment rules (ITRs)

Subgroup Identification

Random forests

Treatment effect heterogeneity

Clinical Trial 

Speaker

Chenguang Zhang, Merck

Co-Author(s)

Duo Yu
han feng
Michael Kane, MD Anderson Cancer Center
Brian Hobbs, University of Texas

Interval Estimation, Hypothesis Testing, and Power Analysis for Fβ scores

Machine learning and artificial intelligence are increasingly applied to medical diagnostics and clinical decision-making. The F1 score and its generalized form, the Fβ score, are commonly used to evaluate diagnostic performance because they effectively balance precision and sensitivity, particularly in the presence of class imbalance. Despite their popularity, rigorous statistical inference and power analysis for these scores remain limited. In this study, we develop methods for interval estimation, hypothesis testing, and power and sample size calculation for both single and comparative F1 and Fβ scores. Extensive simulations demonstrate the accuracy and robustness of the proposed methods. We further showcase their practical utility through real-world biomedical classification applications. These methods enable principled evaluation and comparison of classifiers using F1 and Fβ scores, providing reliable uncertainty quantification and informed sample size planning. Finally, we extend the proposed framework from independent classifier comparisons to settings with correlated classifiers. 

Keywords

F1 score

Class imbalance

Precision

Sensitivity

Sample size calculation 

Speaker

Chih-Yuan Hsu, Vanderbilt University Medical Center

Co-Author(s)

Qi Liu, Vanderbilt University
Yu Shyr, Vanderbilt University Medical Center

Modeling Treatment Switching in Clinical Trials using Machine Learning with Dynamic Clinical Inputs

Traditionally, clinical trial designs treat patients switching from one treatment arm to another as a random process. However, evidence from completed studies shows that crossover occurs in a clinically driven manner and is related to clinical factors, such as disease stage, ECOG performance status and adverse events. The decision to switch treatment arms may also be driven by ethical and safety reasons. We propose using machine learning models to study patterns in treatment switching and build predictive models for treatment switching based on individualized, time varying information such as baseline characteristics, adverse events and changes in disease progression reported throughout follow up. Our proposed framework provides a data driven approach to quantify non random treatment switching patterns. This has potential implications for improved study design, sensitivity analyses, and interpretation of treatment effects in the presence of post randomization treatment switching. 

Keywords

treatment switching

clinical trials

machine learning 

Speaker

Lingli Yang, Takeda

Co-Author(s)

Xuzhi Wang, Takeda Pharmaceutical Company Limited
Deepak Nag Ayyala, Takeda Pharmaceuticals

Zetyra: A Validated, Regulatory-Aligned Calculator Suite for Adaptive and Bayesian Clinical Trial Design

AI-assisted tools are accelerating clinical trial design, but validation has not kept pace. Many tools lack formal validation against reference standards, produce non-reproducible outputs, and fail to capture methodological interactions in composed adaptive designs. These limitations become critical at regulatory submission, where operating characteristics must be defensible under both frequentist and Bayesian frameworks.
Zetyra is a browser-based statistical calculator suite designed to address this validation gap through transparent, reproducible, and regulator-aligned implementations across frequentist, Bayesian, adaptive, and master-protocol design modules. Methods are validated against established reference engines, including gsDesign, RBesT, pwr, and scipy, and validation is enforced through automated tests open to independent verification.
A central contribution is pipeline-level evaluation of composed adaptive designs, where Bayesian monitoring, sample size re-estimation (SSR), and response-adaptive randomization (RAR) update using overlapping interim information. Through simulation, we show that component-level validation is insufficient: under a mild prior-data conflict, Type I error reaches 0.0771 (+208% above nominal α = 0.025) even when individual components pass validation. A negative-control analysis moving SSR to an earlier interim shows negligible attenuation, identifying prior-data conflict rather than SSR timing as the dominant driver. In mechanism-isolation analysis, prior-data conflict and time trends interact super-additively, with the interaction estimate's 95% simulation interval excluding zero.
This talk presents Zetyra's validation framework and demonstrates how composed-design evaluation surfaces operating-characteristic failures that conventional component-wise validation can miss, aligning with the FDA's January 2026 draft guidance on Bayesian methodology in clinical trials. 

Keywords

Adaptive clinical trial design

Bayesian methods

Group sequential design

Sample size re-estimation

Response-adaptive randomization

Validation and reproducibility 

Speaker

Lu Qian, Zetyra

What if We've Been Looking at the Wrong Data? Reimagining Clinical Trial Success Prediction using AI

Background: AI-based clinical trial prediction models (HINT, SPOT) achieved ROC-AUC of 0.65, but complex architectures combining knowledge graphs and Transformers haven't yielded substantial gains, while low explainability limits adoption.

Objective: We hypothesize limitations stem from data quality rather than model architecture, proposing an AI-driven data enrichment approach.

Methods: We audited public benchmarks (TOP, CTO, TrialPanorama) for completeness and label reliability. A feasibility study on 300+ trials evaluated automated annotation using LLMs. We developed an enrichment strategy integrating pharmacokinetic data, preclinical biomarkers, and inter-phase links into an explainable ML model.

Results: Manual annotation of 500 trials revealed labeling errors: 45% (CTO), 9.6% (TrialPanorama), 8.3% (TOP). Automated annotation showed encouraging performance. Missing biological variables significantly limit current models.

Conclusions: This work challenges the paradigm favoring algorithmic complexity. Rigorous data curation and AI-enhanced enrichment could unlock predictive gains with implications for higher probability of success in drug development. 

Keywords

Clinical Trials

Machine Learning

Prediction

Data 

Speaker

Wenting Wang

Co-Author(s)

Leo Fournier, 3University of Montpellier, Montpellier, France
Juan martinez, Cytel, London, UK
Tarun Nathani, Sanofi
Nils Ternes, Sanofi
Christelle Reynes, University of Montpellier, Montpellier, France
Krishna Bellamkonda, Sanofi