pipette
ESEspañol

Errors, Hallucinations, and Clinical Impact of General-Purpose Multimodal Large Language Models in Histopathology

K. Lami, S. Agarwal, A. Asaturova, S. Balci, A. Harahap, H. Kang, J. Kim, M. Komuta, T. Laohawetwanit, S. Menon, J. Munkhdelger, H. H. N. Pham, D. G. Pinto, A. Sahay, S. Satturwar, K. Seki, M. Sughayer, Y. Tachibana, I. Turkmen, Q. Wu, E. Ullah, A. Bychkov, J. Fukuoka

PreprintReal-world use

In the authors' words

Background: General-purpose large language models (LLMs) are increasingly evaluated in diagnostic pathology, but prior studies have largely emphasized diagnostic accuracy rather than how models fail. We evaluated four LLMs for diagnostic performance, pathology-relevant errors and hallucinations, their burden, and potential clinical impact across multi-organ pathology cases. Design: In this retrospective multicenter study, 153 pathology cases from two institutions spanning 20 organs were evaluated using ChatGPT-5.3 (LLM1), Gemini 3 (LLM2), Grok 4.20 (LLM3), and Claude Opus 4.6 (LLM4). Each LLM received multi-magnification histologic images with clinical context and generated a microscopic description and diagnosis. No data splitting, training, or fine-tuning was performed. Twenty-one pathologists assessed 612 outputs for diagnostic correctness, error and hallucination type and burden, clinical impact (0-4), and overall performance (1-5). Results: Strict diagnostic accuracy was 48.9% overall (60.1% including partially correct diagnoses) and ranged from 36.6% to 58.2% across LLMs. Errors occurred in 82.0% of outputs and hallucinations in 76.6%; 90.2% contained at least one error or hallucination. Misinterpretation was the most frequent error (72.1%), while fabricated histologic features were the dominant hallucination type (74.5%). LLM2 had significantly lower misinterpretation rates than the other three LLMs, while LLM3 had significantly higher fabricated-feature hallucination rates than LLM1 and LLM2. Errors were more frequent in incorrect than correct diagnoses (97.5% vs 66.2%), as were hallucinations (95.1% vs 59.9%; both p < 0.001). Every strictly incorrect diagnosis contained an error and/or hallucination, while 79.9% of strictly correct diagnoses also contained at least one. In multivariable analysis, LLM3 was independently associated with higher odds of an incorrect diagnosis. Diagnostic correctness was the strongest determinant of high clinical impact; each one-point increase in hallucination burden increased the odds by 63% (adjusted OR, 1.63). Conclusion: Diagnostic accuracy alone substantially underestimates the safety limitations of general-purpose LLMs in histopathology. Errors and hallucinations were common even when the final diagnosis was correct, and their burden was independently associated with clinical impact. Evaluation frameworks should therefore assess both diagnostic correctness and the reliability and potential consequences of accompanying generated content.

Main resultThe abstract does not state a limitation.

Appeared: Friday, September 25. medRxiv. Preprint, not yet peer-reviewed.

DOI: 10.64898/2026.09.18.26363369