AI 中文总结
本研究通过9,312个成对评判实验,发现LLM评审结果受评审者所属模型系列影响,同系列偏好提升3.4-8.4个百分点,并提出了修正估计量以分离该效应。
AI 中文摘要
评审者是谁会影响LLM作为评审的结果,但衡量这种影响而不与候选质量混淆是困难的。我们研究了四个开放权重系列(Llama 3.1、Qwen 2.5、Gemma 2和Yi 1.5),采用完全交叉的成对设计,共包含9,312个评判。一个常见的按系列统计量与候选质量强烈混淆,并与Bradley-Terry能力在r=0.95处相关。我们推导出一个修正估计量,该估计量固定候选系列并比较评审者。所有四个系列随后都显示出正的同系列提升(3.4-8.4个百分点),全局FPS为0.067(95%置信区间[0.053, 0.084],置换检验p=0.0002)。该效应在基于面板的质量控制、独立的人类共识锚点和float16评判复制下仍然存在。评审者侧似然与效应密切相关:增加似然优势将受控系数降低61%,我们将其视为描述性衰减而非因果中介。位置是另一个独立的失败模式。在整个评审团中,55.4%的AB/BA对发生反转,超过50%的反转率与简单的独立内容噪声模型不相容。相对于系列平衡的参考,评审团组成改变了18.5%的成对结果。已准备完整的可复现性档案供公开发布。
英文摘要
Who the judge is can affect an LLM-as-judge result, but measuring that effect without confusing it with candidate quality is difficult. We study four open-weight families (Llama 3.1, Qwen 2.5, Gemma 2, and Yi 1.5) in a fully crossed pairwise design with 9,312 judgments. A common per-family statistic is strongly confounded with candidate quality and correlates with Bradley-Terry ability at r = 0.95. We derive a corrected estimator that holds the candidate family fixed and compares judges. All four families then show a positive same-family lift (3.4-8.4 percentage points), with global FPS 0.067 (95% CI [0.053, 0.084], permutation p = 0.0002). The effect remains under panel-based quality controls, an independent human-consensus anchor, and a float16 judging replication. Judge-side likelihood is closely related to the effect: adding likelihood advantage reduces the controlled coefficient by 61%, which we treat as descriptive attenuation rather than causal mediation. Position is a separate failure mode. Across the panel, 55.4% of AB/BA pairs reverse, and reversal above 50% is incompatible with a simple independent content-noise model. Relative to a family-balanced reference, panel composition changes 18.5% of pairwise outcomes. A complete reproducibility archive has been prepared for public release.
Comments11 pages, 3 figures, 4 tables