37: Ontology-Aware Evaluation of AI and LLMs for Human Phenotype Ontology Annotation from Clinical Text

Yuyan Yi Speaker
National Institute of Allergy and Infectious Diseases
 
Rachel Waymack Co-Author
National Institute of Allergy and Infectious Diseases
 
Daniel Veltri Co-Author
National Institute of Allergy and Infectious Diseases
 
Monday, Aug 3: 2:00 PM - 3:50 PM
2691 
Contributed Posters 
Thomas M. Menino Convention & Exhibition Center 
Automated annotation of clinical narratives with Human Phenotype Ontology (HPO) terms enables downstream statistical analysis in precision medicine and rare disease research. Advances in large language models (LLMs) have expanded phenotype extraction capabilities, yet their comparative behavior under ontology-aware evaluation remains poorly characterized.

We present a systematic comparison of HPO annotation methods spanning rule-based systems, neural encoder models, and generative LLM-based approaches. Experiments use multiple publicly available and expert-annotated datasets varying in text length and annotation density.

Beyond exact-match metrics, we incorporate HPO hierarchy-aware evaluation and examine the effects of text segmentation, annotation filtering, and hierarchy-based tuning. Results indicate that performance is sensitive to both data characteristics and evaluation design, with methods exhibiting distinct trade-offs in coverage, specificity, and robustness. These findings highlight the importance of ontology-aware evaluation and motivate further methodological refinement.

Keywords

Human Phenotype Ontology

Clinical text analysis

Ontology-aware evaluation

Large language models

Biomedical NLP

Statistical comparison 

Main Sponsor

Section on Text Analysis