pipette
ENEnglish

Extracting smoking history from clinical notes for lung cancer screening decision support: comparing a structured-judgment model with general-purpose large language models

A. Wright, S. Liu, A. P. Wright

PreprintUso en el mundo real

En palabras de los autores

Objective: To compare the accuracy, cost, and speed of a low-cost, non-generative structured-judgment model and four general-purpose large language models (LLMs) for extracting smoking status, pack-years, and quit date from clinical notes to support lung cancer screening (LCS) clinical decision support (CDS). Materials and Methods: We built a synthetic, shareable benchmark of 3,000 outpatient notes in three conditions: 1,000 template-generated notes, 1,000 realistic, "messy" notes written by an LLM from structured facts, and 1,000 "messy" notes that required complex arithmetic to determine pack-years and quit dates. Ground truth reference labels were programmatically generated before each note was created. We compared TypeSafe Jev 1.13 with Claude Haiku 4.5, Claude Sonnet 5, GPT-6 Luna and GPT-6 Sol, each using an identical structured output schema. We determined United States Preventive Services Task Force (USPSTF) 2021 and American Cancer Society (ACS) 2023 eligibility from each system's output in code. Results: Jev made the correct eligibility decision for 99.1%, 99.9%, and 94.4% of notes in the three conditions, compared with 97.8%, 98.4%, and 98.1% for Haiku, 100.0%, 99.6%, and 99.8% for Sonnet, 99.7%, 98.8%, and 98.5% for Luna, and 100.0%, 99.8%, and 99.6% for Sol. In the complex condition, Jev produced 27 false positive and 14 false negative screening flags per 1,000 notes, the most of any system. Luna cost 0.21 per 1,000 notes and Jev 0.64, compared with 4.22 for Sol, 4.43 for Haiku, and 6.61 for Sonnet. Jev was the fastest system (median 0.45 to 1.21 seconds per note, compared with 1.71 to 3.50 seconds for the others). Discussion: All models tested had high accuracy. The structured-judgment model had the lowest latency, which could allow for synchronous use in clinical decision support systems rather than batched, cached invocation. The cost of the structured-judgment model was low, but the newest cost-efficient LLM, Luna, was less expensive and more accurate on the complex notes, though its latency was higher. Accuracy, latency, and cost are all moving quickly as new models are released. Conclusion: Fast, low-cost language models could make real-time, note-based CDS practical, and current models are highly accurate. Prices and capabilities are changing rapidly, so continuous benchmarking and system updates are essential.

Resultado principalLimitación que admiten los autores

Apareció: domingo, 27 de septiembre. medRxiv. Preprint, todavía sin revisión por pares.

DOI: 10.64898/2026.09.24.26363906