pipette
ENEnglish

Leveraging Large Language Models for Colorectal Cancer Symptom Extraction from MIMIC-IV Clinical Notes

Y. Lee, I. Dinov, X. Hu, Y. Jiang

PreprintUso en el mundo real

En palabras de los autores

BackgroundMuch of the symptom burden in colorectal cancer (CRC) patients is documented in unstructured discharge-note narrative, and manual extraction is not scalable. Whether large language models (LLMs) outperform rule-based and named entity recognition (NER) methods has not been rigorously benchmarked. ObjectiveTo benchmark rule-based, NER, and zero-shot LLM methods for extracting 46 cancer-related symptoms from CRC discharge notes against an adjudicated ground truth. MethodsWe analyzed 2,704 discharge notes from CRC patients in MIMIC-IV. A 46-symptom target list was built from the Memorial Symptom Assessment Scale and the EORTC QLQ-CR29. Four approaches -- dictionary-based rule matching, pretrained clinical NER, and zero-shot Claude Haiku and Gemini 3.5 Flash -- plus two hybrid variants (LLM output with post-hoc rule-based negation filtering) were evaluated against a 200-note gold standard adjudicated by two raters (pooled kappa=0.71, macro kappa=0.49), using Macro/Micro F1, precision, and recall. ResultsGemini 3.5 Flash performed best (Macro F1=0.70, Micro F1=0.86, Macro Precision=0.74), followed by Claude Haiku (Macro F1=0.63, Macro Recall=0.71); both substantially outperformed rule-based (Macro F1=0.44) and NER (Macro F1=0.38) methods. Post-hoc negation filtering paradoxically degraded LLM performance (Gemini+Hybrid Macro F1=0.58; Claude+Hybrid Macro F1=0.54) by overriding correct predictions through rigid, fixed-window matching. ConclusionsZero-shot LLMs substantially outperform rule-based and NER approaches for CRC symptom extraction; post-hoc negation correction should not be applied to LLM outputs without syntactic scope validation. Implications for Practice: Zero-shot LLM extraction offers a scalable, accurate alternative to manual chart review and traditional NLP pipelines for oncology symptom surveillance, without institution-specific rule development or model training.

Resultado principalEl resumen no menciona limitaciones.

Apareció: jueves, 24 de septiembre. medRxiv. Preprint, todavía sin revisión por pares.

DOI: 10.64898/2026.09.15.26362961