LLM-Assisted Rare Disease Knowledge Curation Across the Knowledge Lifecycle Using Expert Resources and Real-World EHR Data

Xiaoyi Chen Speaker
 
Tuesday, Aug 4: 11:15 AM - 11:35 AM
Topic-Contributed Paper Session 
Thomas M. Menino Convention & Exhibition Center 
Rare disease knowledge evolves from clinical observations into expert-curated knowledge resources and must be continuously re-evaluated as new real-world evidence becomes available. Effective use of large language models (LLMs) and EHR data for rare disease research therefore requires high-quality knowledge resources beyond generic knowledge bases (KBs) such as OMIM and Orphanet, which are often incomplete, inconsistent, and misaligned with the clinical language used in patient notes. Building and maintaining such resources requires substantial expert effort for knowledge alignment, organization, and disease-specific curation. We investigate how LLM can support this process and how their performance can be rigorously evaluated when the expert reference standard is itself a purpose-built, single-annotator resource.
Using four inherited metabolic diseases (IMDs) spanning the major groups of IMDs and characterized by substantial diagnostic delay, we integrate three complementary knowledge sources: generic knowledge bases, expert-designed diagnostic fact sheets developed by 67 IMD experts, and real-world EHR cohorts from Necker Children's Hospital (Paris, France). We evaluate LLM performance at three stages: (1) aligning phenotype annotations across KBs; (2) aligning expert fact sheets with generic KBs and (3) reproducing physician curation on
1,886 clinical concepts extracted from patient notes for normalization, categorization and diagnostic relevance curation.
Performance is assessed using Cohen's and Fleiss' κ, bootstrap confidence intervals, McNemar's tests, and standard classification metrics.
Across the four IMDs, LLM-assisted phenotype alignment showed only 24.3% exact phenotype agreement between OMIM and Orphanet, while ontology-based semantic relationships explained only 30.4% of the discordant phenotype pairs. LLM-assisted alignment further identified
10 of 104 (9.6%) expert-defined diagnostic phenotypes were absent from both generic KBs. For EHR-derived concepts, the LLM matched physician organ-system categorization for 80.6% of concepts versus 74.9% for an ontology-only approach (p < 0.01), while successfully categorizing the 34% of concepts that could not be mapped by the ontology. For disease-specific relevance curation, a clinical reasoning-grounded LLM reproduced physician keep/drop decisions substantially better than a KB-membership strategy (κ = 0.59 vs. 0.38) and recovered 46% versus 5% of diagnostically relevant phenotypes absent from both generic KBs and expert fact sheets.
These results demonstrate that LLMs are most valuable where existing knowledge resources are incomplete and provide a statistically grounded framework for evaluating and refining rare disease knowledge using complementary evidence from expert resources and real-world EHR data.