arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

我们能信任LLM裁判吗:能力相关偏见与多裁判集成用于偏见校准的研究

Can We Trust LLM Judges: A Study of Capability-Dependent Biases and Multi-Judge Ensemble for Bias Calibration

Gemma Zhang, Prachi Badarayani, Asmi Kumar, Sadid Hasan, Sulaiman Vesal

arXiv 2609.12002首次发表:更新:

发表机构

Microsoft; Massachusetts Institute of Technology(微软; 麻省理工学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究揭示LLM裁判存在能力相关偏见,提出基于分歧估计的加权多数投票集成方法,无需标签即可校准偏见,显著提升评估准确性与公平性。

AI 中文摘要

大型语言模型(LLM)越来越多地被用作模型训练和评估的自动裁判,然而单个裁判会表现出系统性偏见,从而削弱其可靠性。以往的研究大多关注成对LLM裁判场景中的偏见;在本文中,我们聚焦于绝对评分任务,这更贴近实际应用场景。在四个基准和六个模型(共36个裁判-受试者组合)上,我们表明模型的任务准确率能强预测其裁判准确率(大多数模型上Pearson相关系数r ≥ 0.90),并反向预测其方向性偏见(r ≤ -0.83),但仅凭准确率并不能确保公平评估:能力更强的受试者模型始终会从所有裁判那里获得更宽松的评判(r ≥ 0.83)。为解决这一问题,我们提出了校准加权多数投票(WMV),这是一种集成评估方法,通过基于在线估计的假阳性率和假阴性率对多个LLM裁判进行加权聚合。我们引入了一种基于分歧的估计器,该估计器仅从裁判间的一致模式中推导出这些错误率,无需任何真实标签或任务元数据。在任务分布变化的模拟实验中,我们的无标签WMV与具有完美错误率知识的神谕(oracle)的平均差距在0.5个百分点以内,优于单个裁判和未加权多数投票。这些结果表明,原则性的多裁判校准可以在无需标签数据的情况下同时提高准确率并纠正系统性宽松偏见,为随着模型能力提升实现可靠的自动化评估提供了一条可扩展的路径。

英文摘要

LLMs are increasingly used as automated judges for model training and evaluation, yet individual judges exhibit systematic biases that undermine reliability. Much of prior work has studied biases in pairwise LLM-as-a-judge settings; in this paper, we focus on absolute scoring tasks, which mirror more realistic use cases. Across four benchmarks and six models (36 judge-examinee pairs), we show that a model's task accuracy strongly predicts its judging accuracy (Pearson $r \geq 0.90$ on most models) and inversely predicts its directional bias ($r \leq -0.83$), but that accuracy alone does not ensure fair evaluation: more capable examinee models consistently receive more lenient judgments from all judges ($r \geq 0.83$). To address this, we propose calibrated weighted majority voting (WMV), an ensemble evaluation method that aggregates multiple LLM judges weighted by online estimates of their false-positive and false-negative rates. We introduce a disagreement-based estimator that derives these error rates purely from inter-judge agreement patterns, requiring no ground-truth labels or task metadata. In a simulated experiment with shifting task distributions, our label-free WMV tracks an oracle with perfect error-rate knowledge to within 0.5 percentage points on average, outperforming both individual judges and unweighted majority voting. These results demonstrate that principled multi-judge calibration can simultaneously improve accuracy and correct for systematic leniency without requiring labeled data, offering a scalable path to reliable automated evaluation as model capabilities increase.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑