LLM-enabled Natural History Study Analysis to Support Rare Disease Research
En palabras de los autores
BackgroundRare diseases affect an estimated 300 million people worldwide, yet the research needed to guide diagnosis and treatment is often fragmented across multiple unstructured literature sources. Natural history studies (NHS) are a key source of this evidence, but manually extracting structured information from NHS publications can be tedious and does not scale. MethodsWe have developed a proof-of-concept for an information extraction pipeline testing three open-source large-language models (LLMs) - Athena-v3-AWQ, Googles Gemma3-27B, and Metas Llama-3.1-70B-Instruct, to extract key NHS characteristics from PubMed abstracts curated from a Chan Zuckerberg Initiative disease research state model corpus (302 gold-standard and 8,338 full-corpus abstracts), and compared the models on efficiency, extraction completeness, and expert-rated accuracy. ResultsAll three models processed abstracts with success rates exceeding 99%. However, Gemma achieved the best overall performance, with the highest expert-rated accuracy (68.0% of outputs rated "good" vs. 36.0% for Llama and 10.0% for Athena) and the fastest runtime on the full corpus ([~]16 minutes for 3,547 abstracts), despite Llama scoring higher on the automated Token F1 metric (0.874 vs. 0.723), highlighting a divergence between automated and human evaluation. Athenas lower performance was largely attributable to verbatim copying rather than synthesis of extracted content. ConclusionsThese findings illustrate how locally deployed open-source LLMs can extract structured NHS characteristics at scale, thus supporting their use to accelerate evidence synthesis in rare disease research.
Apareció: jueves, 24 de septiembre. medRxiv. Preprint, todavía sin revisión por pares.