pipette
ESEspañol

An LLM-Based Auditing Framework for Targeted Quality Assurance in Cancer Registries

M. Seifaddini, M. Beheshti, J. Steffens, D. Carnagey, M. Esebua, M. Wakefield, I. Zachary

PreprintReal-world use

In the authors' words

Cancer registries are essential for population-level surveillance, epidemiologic research, and public health planning, but their accuracy depends on human review of key variables documented in unstructured clinical text. This is time-consuming, and current quality assurance (QA) typically relies on reviewing fewer than 10% of records at random, which can miss facility- or time-specific errors. We developed a large language model (LLM)-based framework for targeted cancer registry QA. The framework uses LLMs to extract biomarker values from clinical text, compares them with registry entries, and flags disagreements for expert review. Six open-source LLMs were benchmarked locally in a privacy-preserving environment, and the best model per site and report type was applied to breast and prostate cancer, two of the most common cancers in the United States. For the four breast cancer biomarkers ER, PR, HER2, and Ki-67, the disagreement rate against quality-assured registry values in 1,852 breast cancer cases was lowest for ER at 1.67% and highest for Ki-67 at 9.18%. In 920 prostate cancer cases, the disagreement rate with the deployed models stayed below 8% for every categorical Gleason variable, while the exact-value PSA metric reached 17.28%. PSA therefore generated the largest volume of candidate discrepancies, though individual PSA flags require more corroboration than categorical flags. Applied to a full year of records (11,740 breast and 6,792 prostate cases), the framework flagged disagreements and assigned each a probable cause, grounded in NAACCR coding rules, to prioritize expert review. Extracted values are linked to their supporting text, and model, prompt, and pipeline versions are logged for reproducibility. By directing review toward records most likely to contain discrepancies, the framework offers an efficient, auditable alternative to random QA sampling for cancer registries.

Main resultLimitation the authors admit

Appeared: Thursday, September 24. medRxiv. Preprint, not yet peer-reviewed.

DOI: 10.64898/2026.09.22.26363595