What if We've Been Looking at the Wrong Data? Reimagining Clinical Trial Success Prediction using AI

Wenting Wang Speaker
 
Leo Fournier Co-Author
3University of Montpellier, Montpellier, France
 
Juan martinez Co-Author
Cytel, London, UK
 
Tarun Nathani Co-Author
Sanofi
 
Nils Ternes Co-Author
Sanofi
 
Christelle Reynes Co-Author
University of Montpellier, Montpellier, France
 
Krishna Bellamkonda Co-Author
Sanofi
 
Monday, Aug 3: 12:05 PM - 12:20 PM
2963 
Contributed Papers 
Thomas M. Menino Convention & Exhibition Center 
Background: AI-based clinical trial prediction models (HINT, SPOT) achieved ROC-AUC of 0.65, but complex architectures combining knowledge graphs and Transformers haven't yielded substantial gains, while low explainability limits adoption.

Objective: We hypothesize limitations stem from data quality rather than model architecture, proposing an AI-driven data enrichment approach.

Methods: We audited public benchmarks (TOP, CTO, TrialPanorama) for completeness and label reliability. A feasibility study on 300+ trials evaluated automated annotation using LLMs. We developed an enrichment strategy integrating pharmacokinetic data, preclinical biomarkers, and inter-phase links into an explainable ML model.

Results: Manual annotation of 500 trials revealed labeling errors: 45% (CTO), 9.6% (TrialPanorama), 8.3% (TOP). Automated annotation showed encouraging performance. Missing biological variables significantly limit current models.

Conclusions: This work challenges the paradigm favoring algorithmic complexity. Rigorous data curation and AI-enhanced enrichment could unlock predictive gains with implications for higher probability of success in drug development.

Keywords

Clinical Trials

Machine Learning

Prediction

Data 

Main Sponsor

Biopharmaceutical Section