pipette
ENEnglish

Calibrating LLM Judges for Human and AI Conversations

Maike Z\"ufle, Patr\'icia Schmidtov\'a, Vil\'em Zouhar, Shree Harsha Bokkahalli Satish, Erica Cooper, Shobhit Banga, Vaibhav Nalawade, Manmeet Kaur, Jan Niehues, Markus M\"uller, Ond\v{r}ej Klejch

Preprint

En palabras de los autores

Measuring how successful a conversation is remains difficult, even for humans judging spoken dialogue. We evaluate state-of-the-art LLMs as pointwise and pairwise judges of conversational success on CANDOR, finding pointwise scoring correlates moderately with human ratings, while pairwise comparison suffers from long transcripts and positional bias. Since this leaves judge scores incomparable across models, we propose a small anchor set and a calibration function that calibrates any judge onto a shared, interpretable scale. We further release the Voice Arena Goal Dataset (VA), 200 task-oriented human-AI and human-agent conversations with pairwise annotations, revealing a substantial gap between current judges and human-level discrimination. Using VA, we test whether CANDOR-fitted calibration transfers to human-AI conversations, finding it brings judges onto a shared scale despite never observing VA during fitting.

Resultado principalLimitación que admiten los autores

Apareció: viernes, 25 de septiembre. arXiv. Preprint, todavía sin revisión por pares.