pipette
ESEspañol

RoB2-Evaluator: Development and Technical Evaluation of a Rule-Constrained Large Language Model System for Cochrane Risk of Bias 2 Assessment

J. T. Joseph, R. Vishwanath, J. V. Sisirkumar

PreprintReal-world use

In the authors' words

BackgroundCochrane Risk of Bias 2 (RoB 2) assessment is methodologically demanding and resource intensive. Large language models may assist this process but can produce variable judgments. AimTo develop a rule-constrained LLM system for RoB 2 assessment and evaluate its temporal stability and agreement with human reviewers. MethodsRoB2-Evaluator was developed using GPT-5.2 with fixed RoB 2 guidance and a locked instruction framework. Seven RCTs were used for calibration before version freezing. The frozen system was evaluated on 10 separate RCTs, with repeat assessment after one week and comparison against consensus judgments from two independent human reviewers. ResultsTemporal agreement was 90.0% (weighted {kappa} = 0.93), and agreement with human consensus was also 90.0% (weighted {kappa} = 0.93). Methodological rationales were stable in 94.3% of repeated comparisons. Full AI-human conceptual alignment occurred in 85.7%. ConclusionRoB2-Evaluator showed encouraging post-freeze stability, human agreement, and rationale consistency, supporting larger independent evaluations.

Main resultLimitation the authors admit

Appeared: Thursday, September 24. medRxiv. Preprint, not yet peer-reviewed.

DOI: 10.64898/2026.09.22.26363693