pipette
ENEnglish

An auditable evidence compiler for large language model-assisted systematic reviews

C. Yin, Z. Jing, Z. Zhang

Preprint

En palabras de los autores

BackgroundLarge language models (LLMs) can support systematic reviews, but accurate individual outputs do not establish whether the final synthesis preserves the clinical question, accounts for statistical dependence and incorporates corrections. ObjectiveTo develop and evaluate a framework linking LLM-assisted evidence processing to a versioned, auditable release of synthesis outputs. MethodsWe used LLM agents to interpret sources and extract data. Deterministic code enforced statistical rules; investigators resolved material ambiguities and authorized release. We specified ten release properties covering evidence identities, statistical contributions and propagation of corrections. We retrospectively evaluated six integrity domains and historical failure events in one registered prognostic review, without an external comparator or held-out domain. ResultsFifty distinct root-cause events were documented, including 12 that had changed a pooled result before correction. Forty-six were resolved, and four remained disclosed limitations. The corpus comprised 454 reports, 445 studies, 441 cohort entities and 421 dependence clusters. Forty-one of 49 registered analyses were fitted, and eight retained explicit non-fitted states. All 39 source records across five principal analysis families reached a terminal source state. Two implementations within the project agreed across 1,217 numerical comparisons. All 94 file comparisons between release and publication packages were byte-identical. Two reviewers confirmed 39 principal records after seeing the same recommendations. ConclusionsThis case provides a framework for inspecting synthesized evidence together with its provenance, statistical meaning and correction history. Comparative validity, generalizability and benefit in patient-centred care require independent evaluation. HighlightsO_ST_ABSWhat is already knownC_ST_ABSLLM-assisted systems support screening, extraction, synthesis and review updating. Prior work includes reviewable evidence packages, source verification and audit-guided retrieval. Evaluations include task benchmarks, review replication and effects of corrected inputs on pooled results. What is newOur implementation links source judgments, typed evidence identities and analysis-specific contributions to the consequences of corrections and the status of released artifacts. We evaluate its conformance and limits in a production-scale systematic review. Fifty root-cause events include 12 that had changed a pooled result before correction. Potential impact for Research Synthesis Methods readersReaders can inspect links among clinical questions, source judgments, statistical contributions and released artifacts. Assumptions and limitations can inform evidence reuse; documented failures identify controls to test in other workflows. Patient-centred application still requires assessment of applicability, preferences and clinical outcomes.

Resultado principalLimitación que admiten los autores

Apareció: miércoles, 23 de septiembre. medRxiv. Preprint, todavía sin revisión por pares.

DOI: 10.64898/2026.09.21.26363538