pipette
ESEspañol

Accounting for Bias Enables Sustainable LLM Evaluation

Harshita Katoch, David Antony Selby, Gerrit Gro{\ss}mann, Sebastian Vollmer

PreprintBold claims, read critically

In the authors' words

LLM-as-a-judge has become the de facto standard for scalable, subjective evaluation, yet current leaderboards compensate for systematic measurement bias by running ever more comparisons, an approach that is both statistically unsound and computationally wasteful. The root cause is an incomplete measurement model, treating LLM judges as neutral, interchangeable instruments ignores documented biases like position bias, verbosity bias, judge severity, and self-enhancement, that no volume of additional data can eliminate. We propose a unified latent variable framework that jointly models pairwise and ordinal data while explicitly correcting for these confounders, recovering reliable rankings from substantially fewer comparisons. Because fitting this model costs negligible compute relative to a single round of LLM inference, bias correction is not only more statistically rigorous but also a more sustainable approach to trustworthy evaluation.

Main resultThe abstract does not state a limitation.

Appeared: Monday, September 28. arXiv. Preprint, not yet peer-reviewed.

Authors' comment: 8 pages, 2 figures; SuRE'26: Workshop on Sustainability and Resource-Efficiency of Artificial Intelligence at IJCAI-ECAI 2026