arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

谁来做评判很重要:衡量LLM评审团中的家庭条件偏好

Who Judges Matters: Measuring Family-Conditioned Preference in LLM-as-Judge Panels

David Ababio Awuni, Luke E. K. Achenie, Benjamin Tei Partey, Elvis Gyasi Owusu, Nii-Nai Derrick Sowah

arXiv 2609.17857首次发表:更新:

AI 中文总结

本研究通过9,312个成对评判实验,发现LLM评审结果受评审者所属模型系列影响,同系列偏好提升3.4-8.4个百分点,并提出了修正估计量以分离该效应。

AI 中文摘要

评审者是谁会影响LLM作为评审的结果,但衡量这种影响而不与候选质量混淆是困难的。我们研究了四个开放权重系列(Llama 3.1、Qwen 2.5、Gemma 2和Yi 1.5),采用完全交叉的成对设计,共包含9,312个评判。一个常见的按系列统计量与候选质量强烈混淆,并与Bradley-Terry能力在r=0.95处相关。我们推导出一个修正估计量,该估计量固定候选系列并比较评审者。所有四个系列随后都显示出正的同系列提升(3.4-8.4个百分点),全局FPS为0.067(95%置信区间[0.053, 0.084],置换检验p=0.0002)。该效应在基于面板的质量控制、独立的人类共识锚点和float16评判复制下仍然存在。评审者侧似然与效应密切相关:增加似然优势将受控系数降低61%,我们将其视为描述性衰减而非因果中介。位置是另一个独立的失败模式。在整个评审团中,55.4%的AB/BA对发生反转,超过50%的反转率与简单的独立内容噪声模型不相容。相对于系列平衡的参考,评审团组成改变了18.5%的成对结果。已准备完整的可复现性档案供公开发布。

英文摘要

Who the judge is can affect an LLM-as-judge result, but measuring that effect without confusing it with candidate quality is difficult. We study four open-weight families (Llama 3.1, Qwen 2.5, Gemma 2, and Yi 1.5) in a fully crossed pairwise design with 9,312 judgments. A common per-family statistic is strongly confounded with candidate quality and correlates with Bradley-Terry ability at r = 0.95. We derive a corrected estimator that holds the candidate family fixed and compares judges. All four families then show a positive same-family lift (3.4-8.4 percentage points), with global FPS 0.067 (95% CI [0.053, 0.084], permutation p = 0.0002). The effect remains under panel-based quality controls, an independent human-consensus anchor, and a float16 judging replication. Judge-side likelihood is closely related to the effect: adding likelihood advantage reduces the controlled coefficient by 61%, which we treat as descriptive attenuation rather than causal mediation. Position is a separate failure mode. Across the panel, 55.4% of AB/BA pairs reverse, and reversal above 50% is incompatible with a simple independent content-noise model. Relative to a family-balanced reference, panel composition changes 18.5% of pairwise outcomes. A complete reproducibility archive has been prepared for public release.

Comments11 pages, 3 figures, 4 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑