arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.11871cs.LGcs.AIcs.CL

剖析不公平的评判者:基于大语言模型作为评判者的偏差的机械可解释性分析

Inside the Unfair Judge: A Mechanistic Interpretability Account of LLM-as-Judge Bias

Zixiang Xu, Sixian Li, Huaxing Liu, Xiang Wang, Shuai Li, Zirui Song, Xiuying Chen

首次发表
浏览论文内容

中文总结 AI 辅助

研究大语言模型作为评判者的偏差,提出偏差在隐藏状态有表示层面解释。通过七位评判者等实验发现,偏差输入沿特定子空间移动,操纵隐藏状态可控制评分,线性投影能预测评判失败,统一多方面内容。

中文摘要 AI 辅助

现有的关于大语言模型作为评判者评分偏差的研究主要在输入-输出层面进行:扰动输入、测量分数变化并提出提示级别的缓解措施。我们认为,同样的偏差在评判者的隐藏状态中存在一种表示层面的解释,它与输入-输出观点互补,并且在操作上具有实用价值。我们报告了三项发现,涉及七位评判者、七种偏差类型和九个基准。几何学方面:基线评判输入占据一个紧密的激活流形,而有偏差的输入则沿着一个低维的、特定类型的子空间移动,该子空间随深度而锐化,并能被三类估计器一致地恢复。因果控制方面:沿着这个子空间操纵隐藏状态会在两个方向上驱动评分,正向移动在干净输入上重现偏差评分,反向移动在有偏差的输入上恢复基线评分,而匹配范数的随机方向产生的移动则小一个数量级。操作方面:在相同偏差方向特征上的简单线性投影能够预测在三个完全未见过的基准上的评判失败,显著优于基于文本的替代方法。将偏差视为激活几何学,而非输入-输出噪声,在一个单一框架内统一了几何结构、因果控制和操作预测。项目页面可在这个https链接获取。

英文摘要

Existing studies of LLM-as-judge scoring bias work predominantly at the input-output level: they perturb inputs, measure score deltas, and propose prompt-level mitigations. We argue that the same biases admit a representation-level account in the judge's hidden state, complementary to the input-output view and operationally useful in ways it does not afford. We report three findings, across seven judges, seven bias types, and nine benchmarks. Geometry: baseline judging inputs occupy a tight activation manifold while biased inputs are displaced along a low-dimensional, type-specific subspace that sharpens with depth and is recovered consistently by three families of estimators. Causal control: steering hidden states along this subspace drives scoring in both directions, forward shifts reproducing biased scoring on clean inputs and reverse shifts restoring baseline scoring on biased ones, while matched-norm random directions produce shifts an order of magnitude smaller. Operational: a simple linear projection onto the same bias-direction features anticipates judge failures on three entirely unseen benchmarks, substantially outperforming text-based alternatives. Reading bias as activation geometry, rather than as input-output noise, unifies geometric structure, causal control, and operational prediction within a single framework. The project page is available at https://xzx34.github.io/unfair-judge/

发表机构

  • AMAP, Alibaba Group(阿里巴巴集团AMAP实验室)
  • Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学)
  • University of Southern California(南加州大学)
  • University of Michigan, Ann Arbor(密歇根大学安娜堡分校)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑