Validating LLM judges for automated oversight of patient communication
In the authors' words
LLMs are increasingly used to mediate patient communication, yet scalable evaluation of their safety, accuracy, and communication quality remains an open problem. LLM judges have emerged as automated evaluators, but whether they can holistically replicate human expert judgment is unvalidated. Informed consent for clinical trials presents a demanding case for such validation because it requires conveying complex information to lay audiences under ethical and safety constraints. We developed a stakeholder-informed seven-criterion evaluation rubric spanning safety, reliability, and communication quality. Clinician reference ratings showed strong interrater reliability across all criteria. We validated the rubric on the Informed CONsent Benchmark (ICON-Bench) and benchmarked 19 LLM judges across multiple implementation strategies. LLM judges achieved strong clinician agreement for safety screening and factual verification (Spearman{rho} > 0.80) but weaker agreement for communication quality ({rho} < 0.60). Safety-specialized guard models underperformed general-purpose models. Patient advocates rated communication quality lower than both clinicians and LLM judges. These findings support LLM judges for scalable patient communication oversight while demonstrating the need for recalibration to patient-centered evaluation standards.
Appeared: Wednesday, September 23. medRxiv. Preprint, not yet peer-reviewed.