自标签与他标签诱导大语言模型评判器产生双向偏差
Self- and Other-Labels Induce Bidirectional Bias in LLM Judges
浏览论文内容
中文总结 AI 辅助
该研究针对LLM评判器的自偏好混杂问题,通过评估带模型特征的叙事约束选择,发现控制质量后自偏好消失,而自/他标签会双向影响评分,明确了作者归属是评估偏差的独立驱动因素。
中文摘要 AI 辅助
随着大语言模型(LLM)作为评判者的系统日益普及,LLM的自偏好——即倾向于偏好自身输出的特性——引发了人们对评估可靠性的日益担忧。然而,这一特性主要是在生成文本上研究的,在生成文本中,风格特征和响应质量不可避免地混杂在一起。因此,现有的测量方法无法将真正的自偏好与这些混杂因素区分开。我们通过改变评估对象来解决这一问题:不是评判生成文本,而是让10个LLM评估叙事约束选择,这些选择不带有特定模型的风格特征,但保留了可恢复的特定模型特征。我们进行了两项实验,得出了不同的发现。在盲评条件下,一旦控制了选择质量和评估者的严格程度,自偏好基本消失。在四个评分维度中,三个维度上自偏好消失,第四个维度上则发生反转,即评判者将自己的选择评为原创性更低。然而,在质量匹配的条件下,仅自标签和他标签——不提及任何模型名称——就会双向改变分数:LLM评判者会提高自标签选择的分数,降低他标签选择的分数,无论选择的实际来源是什么。我们做出了两项贡献:1)作者归属是评估偏差的一个独立驱动因素;2)无真实值的开放式任务可作为研究LLM评判者行为的受控工具。
英文摘要
As LLM-as-a-judge becomes increasingly widespread, self-preference -- the tendency of a judge to favor its own outputs -- raises growing concerns about evaluation reliability. However, this bias has been studied predominantly on generated text, where stylistic features and response quality are inevitably conflated. As a result, existing measurements cannot separate genuine self-preference from these confounds. We address this limitation by changing the object of evaluation: instead of judging generated text, ten LLMs assess sets of narrative constraints selected from a shared pool, which carry no stylistic fingerprint yet retain a recoverable model-specific signature. Two experiments on this task yield complementary findings. Under blind evaluation, self-preference disappears, with a small effect remaining in the opposite direction once selection quality and judge severity are controlled. Under matched quality, however, self- and other-labels alone -- without naming any model -- shift scores bidirectionally. LLM judges inflate scores for self-labeled selections and deflate those for other-labeled ones regardless of the selection's actual source. We make two contributions: 1) authorship attribution is a distinct driver of evaluation bias, and 2) ground-truth-free tasks can serve as controlled instruments for studying LLM judge behavior.
发表机构
- School of Digital Humanities and Computational Social Sciences, KAIST(韩国科学技术院数字人文与计算社会科学学院)
机构由 AI 辅助整理,请以论文原文为准。