发表机构
University of Stuttgart(斯图加特大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究系统比较了LLM作为裁判时提示设计、评分量表和模型选择的影响,发现模型身份是方差主要来源,设计选择可显著影响宽松度,但总体裁判可靠,为自动化评估提供方法论参考。
AI 中文摘要
研究人员越来越多地使用大型语言模型作为裁判(LLM-as-a-judge)来评估模型输出。然而,目前对于如何设计这些裁判尚无标准。通常,研究人员凭直觉选择提示、评分量表和模型。如果这些选择改变了裁判的裁决,两项研究可能会对相同的事实得出不同的结论。为应对这一风险并为裁判设计提供实证基础,我们在两个任务上评估了多种设计下的10个推理模型:对句子情感和毒性进行标量评分(每类超过500个项目),以及对问答对进行二元准确性分类(n=600)。对于评分任务,尽管裁判与人类真实情况存在显著分歧,但差异的实际大小足以认为大多数裁判是可靠的(在1-7分制上平均绝对偏差为0.11分);毒性裁判甚至优于标准分类器。在准确性分类任务中,裁判的平均准确率也高达96.5%。然而,设计选择可能产生偏移:仅改变评分量表就可能使测量的偏差偏移高达0.93分(评分任务),而虽然准确率水平很少受到影响,但设计选择始终影响裁判的宽松度(分类任务;使用详细提示时宽松度下降28.9个百分点,更换模型时下降高达56.1个百分点)。与直觉相反,较低的推理努力既不影响准确性也不影响宽松度。在两个任务中,模型身份是方差的主要来源。这些发现表明,虽然LLM裁判在总体上值得信赖,但设计选择可能是方差的重要来源。鉴于LLM研究中自动化评估的依赖日益增加,我们打算将这项研究作为设计更稳健和可复现的LLM-as-a-judge流水线的方法论参考。
英文摘要
Researchers increasingly use Large Language Models as judges (LLM-as-a-judge) to evaluate model outputs. Yet there are no standards for how to design these judges. Typically, researchers choose the prompt, rating scale, and model intuitively. If these choices change the judge's verdicts, two studies can reach different conclusions about the same facts. To address this risk and to provide an empirical basis for judge designs, we evaluate 10 reasoning models across multiple designs on two tasks: a scalar rating of sentence sentiment and toxicity (over 500 items per category), as well as a binary accuracy classification of question-answer pairs (n=600). For the rating tasks, despite judges showing significant disagreements with the human ground truth, the practical size of differences is small enough to consider most judges reliable (mean absolute deviation of 0.11 points on a 1 - 7 scale); toxicity judges even outperform standard classifiers. Judges are also highly accurate on average (96.5%) for the accuracy classification task. However, design choices can produce shifts: changing the rating scale alone can shift measured bias by up to 0.93 points (rating task), and while accuracy levels are rarely impacted, design choices consistently impact judge leniency (classification task; leniency drop of 28.9 percentage points when using detailed prompts, and up to 56.1 percentage points when switching models). Counterintuitively, lower reasoning effort affects neither accuracy nor leniency. Across both tasks, model identity is the dominant source of variance. These findings suggest that while LLM judges are broadly trustworthy in aggregate, design choices can be meaningful sources of variance. Given the growing reliance on automated evaluation in LLM research, we intend this study as a methodological reference for designing more robust and replicable LLM-as-a-judge pipelines.