arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.31184cs.AIcs.LGstat.APstat.ME

考虑偏差可实现可持续的LLM评估

Accounting for Bias Enables Sustainable LLM Evaluation

Harshita Katoch, David Antony Selby, Gerrit Großmann, Sebastian Vollmer

首次发表
浏览论文内容

中文总结 AI 辅助

针对LLM评判中的系统性偏差,提出统一潜变量框架联合建模成对与序数数据并显式校正混杂因素,以更少比较恢复可靠排名,实现统计严谨且可持续的评估。

中文摘要 AI 辅助

以LLM为评判者已成为可扩展、主观评估的事实标准,然而当前的排行榜通过进行越来越多的比较来补偿系统性测量偏差,这种方法在统计上不健全且在计算上浪费。根本原因在于测量模型不完整,将LLM评判者视为中立、可互换的工具忽略了已记录的偏差,如位置偏差、冗长偏差、评判者严格性和自我增强,这些偏差无法通过增加数据量来消除。我们提出一个统一的潜变量框架,该框架联合建模成对数据和序数数据,同时显式校正这些混杂因素,从而从显著更少的比较中恢复可靠的排名。由于相对于单轮LLM推理,拟合此模型的计算成本可忽略不计,偏差校正不仅在统计上更严谨,而且也是实现可信评估的更可持续的方法。

英文摘要

LLM-as-a-judge has become the de facto standard for scalable, subjective evaluation, yet current leaderboards compensate for systematic measurement bias by running ever more comparisons, an approach that is both statistically unsound and computationally wasteful. The root cause is an incomplete measurement model, treating LLM judges as neutral, interchangeable instruments ignores documented biases like position bias, verbosity bias, judge severity, and self-enhancement, that no volume of additional data can eliminate. We propose a unified latent variable framework that jointly models pairwise and ordinal data while explicitly correcting for these confounders, recovering reliable rankings from substantially fewer comparisons. Because fitting this model costs negligible compute relative to a single round of LLM inference, bias correction is not only more statistically rigorous but also a more sustainable approach to trustworthy evaluation.

发表机构

  • University of Kaiserslautern–Landau (RPTU)(凯泽斯劳滕-兰道大学)
  • German Research Centre for Artificial Intelligence (DFKI)(德国人工智能研究中心)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑