发表机构
University of Potsdam; National Research Council(波茨坦大学; 国家研究委员会)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
探讨在无参考答案时大语言模型评判器能否可靠评估。通过校准和敏感性两阶段实验,发现其在无参考答案时易高估错误答案,添加参考答案信息会大幅改变评判结果,且与人类判断相符,强调校准LLM评判器的必要性并提供了方法。
AI 中文摘要
大语言模型(LLM)评判器越来越多地用于评估开放式模型的回答,通常是在没有参考答案的情况下。然而,它们能否在这种评估设置中可靠地进行评估?本文通过两阶段流程探讨了这个问题:一是校准实验,评估评判模型对所评估任务的了解;二是敏感性实验,评估提示中参考答案的存在和位置如何影响评判模型的性能。在涵盖三种语言的实验中,我们发现评估的评判模型在没有参考答案时往往对错误答案评价过高,在某些实验设置中,向提示中添加参考答案信息会使评判模型的正误判断翻转多达85%。与部分人工标注的比较表明,这些由参考答案驱动的变化通常与人类判断一致。我们的结果强调了在将LLM评判器可靠地用于无参考设置之前,需要用有参考意识的评估样本对其进行校准,我们的方法为研究人员和从业者对LLM评判器进行此类校准提供了蓝图。
英文摘要
LLM judges are increasingly being used to evaluate open-ended model responses, often in no-reference settings where a ground-truth answer is unavailable. However, can they reliably assess in such evaluation setups? We explore this question in this paper through a two stage pipeline with a) calibration experiments that assess the judge model's knowledge of the task it is evaluating, and b) sensitivity experiments that assess how the judge model's performance is impacted by the presence and positioning of the reference answer in the prompt. Across experiments covering three languages, we show that the judge models we evaluated tend to over-credit incorrect answers in the absence of a reference answer, and adding reference answer information to the prompt flips the judge model's correct/incorrect decisions by as much as 85% in some experimental settings. Comparison with a subset of human annotations shows that these reference-driven changes generally align with human judgments. Our results emphasize the need for calibrating the LLM judges with a sample with reference-aware evaluation before using them in reference-free setups reliably, and our methodology provides a blueprint for researchers and practitioners in doing such calibration of LLM judges for other tasks.
CommentsPreprint