pipette
ESEspañol

Calibrating LLM Judges for Human and AI Conversations

Maike Z\"ufle, Patr\'icia Schmidtov\'a, Vil\'em Zouhar, Shree Harsha Bokkahalli Satish, Erica Cooper, Shobhit Banga, Vaibhav Nalawade, Manmeet Kaur, Jan Niehues, Markus M\"uller, Ond\v{r}ej Klejch

Preprint

In the authors' words

Measuring how successful a conversation is remains difficult, even for humans judging spoken dialogue. We evaluate state-of-the-art LLMs as pointwise and pairwise judges of conversational success on CANDOR, finding pointwise scoring correlates moderately with human ratings, while pairwise comparison suffers from long transcripts and positional bias. Since this leaves judge scores incomparable across models, we propose a small anchor set and a calibration function that calibrates any judge onto a shared, interpretable scale. We further release the Voice Arena Goal Dataset (VA), 200 task-oriented human-AI and human-agent conversations with pairwise annotations, revealing a substantial gap between current judges and human-level discrimination. Using VA, we test whether CANDOR-fitted calibration transfers to human-AI conversations, finding it brings judges onto a shared scale despite never observing VA during fitting.

Main resultLimitation the authors admit

Appeared: Friday, September 25. arXiv. Preprint, not yet peer-reviewed.