Fully Automated Abstraction of Longitudinal Breast Oncology Records with Off-The-Shelf Large Language Models
In the authors' words
Manual chart abstraction is a major bottleneck in clinical research. In oncology, important outcomes are often documented only in clinical notes. We developed an open-source pipeline that, in a HIPAA-compliant setting, uses general-purpose large language models (LLMs) to abstract variables from longitudinal patient records without fine-tuning or hand-selecting input documents. We randomly selected 100 patients from a complex institutional breast cancer cohort and compared four LLMs with oncologist abstraction, genetic testing records, and, for select variables, research coordinator abstraction. The unedited charts contained a median of 3,100 pages of text, 7 lines of therapy, and 6.5 years of follow-up. Across the two best-performing models, GPT-5 and Gemini 2.5 Pro, the pipeline achieved 90%-100% concordance for recurrence, BRCA1/2, hormone receptor, HER2, clinical stage, PIK3CA, and ESR1 status. Anti-cancer drug identification was comparable to that of a second oncologist. Therapy-line reconstruction was less accurate than a second oncologist; all four LLMs outperformed research coordinators on both tasks. In a separate cohort of 97 patients, the unmodified pipeline showed similar performance for recurrence detection. These findings suggest that LLMs may enable scalable abstraction from narrative medical records, but replication and validation across institutions and diseases are needed.
Appeared: Tuesday, September 22. medRxiv. Preprint, not yet peer-reviewed.