AI 中文总结
针对LLM裁判存在噪声和系统性偏差(如偏好冗长或格式好的回答)的问题,提出将裁判过程建模为贝叶斯推断,引入显式的裁判特定偏差协变量,并设计一种Top-k感知的主动获取规则,以在固定比较预算下准确识别Top-k项。
AI 中文摘要
大型语言模型(LLM)越来越多地被用作廉价、可扩展的裁判,它们成对比较候选输出——用于对回答排序、选择模型或筛选论文。然而,LLM裁判既有噪声又存在系统性偏差:它们偏好冗长或格式良好的答案,并表现出位置效应,因此简单地聚合它们的投票得到的是呈现方式的排序,而非真实质量的排序。我们研究在固定比较预算下识别Top-k项的实际目标,并做出两项贡献。首先,我们将裁判过程建模为对潜在质量的贝叶斯推断,并显式包含裁判特定的偏差协变量(冗长度、位置),通过收缩先验进行正则化,使得数据决定给定裁判实际表现出的偏差。其次,我们引入一种Top-k感知的主动获取规则,选择下一个比较以最大程度减少关于Top-k成员身份的不确定性,而非关于完整排序的不确定性。在一个已知真实质量的可控基准上,由16个真实LLM(涵盖开源和专有系列:Llama、Qwen、Phi-4、GPT-4o-mini/5.1/5.5、Gemini、DeepSeek和Claude Haiku/Sonnet/Opus)进行裁判,朴素聚合在存在偏差的裁判上无论预算多少都会停滞在错误的Top-k上,而我们的偏差感知模型能够恢复正确的Top-k;Top-k感知的获取规则达到这一上限所需的比较次数远少于循环赛或全局不确定性(D-最优)规则。偏差是真实存在的,但具有异质性和能力依赖性:廉价和中档裁判表现出强烈的冗长度偏差,我们的模型能够纠正(将召回率从约0.5–0.6提升到0.84–1.0),而我们测试的前沿裁判几乎没有偏差且已经能够准确排序,因此偏差感知建模在那里变化不大。
英文摘要
Large language models (LLMs) are increasingly used as cheap, scalable judges that compare candidate outputs pairwise. Because such judges prefer verbose or well-formatted answers, the natural fix is to add bias covariates to a Bradley--Terry model and estimate the bias away. We show this cannot work as advertised: the quality/bias split is \emph{not identified} by pairwise comparisons, and the failure is exact -- across $48$ real judge-pools the profile likelihood over the coefficient is flat to $\mathbf{0.0000}$ \textbf{nats}, and scaling the comparisons $26\times$ buys none. A ``debiased'' score is selected by the prior, not recovered from data. Our contribution is accordingly not a better estimator but a characterization of \emph{when prior-based correction is justified}, plus designs that supply the missing information when it is not. The assumption the prior encodes -- quality is a priori uncorrelated with the covariate -- pays only while $\mathrm{corr}(θ,x)$ stays below a crossing point (configuration-dependent, $0.22$--$0.60$), which is what makes the same model help on LLMBar and hurt on SummEval and Nectar. We give two escapes: a \textbf{trusted-anchor gate} that decides per (judge, covariate, task) (no false enables in $6{,}000$ decisions at $K\ge10$ anchors, a rate our sample bounds at $\le6\%$), and a \textbf{paired rendering design}. Across fifteen real LLM judges bias is heterogeneous and capability-dependent: correction improves \topk{} recall by $0.20$--$0.32$ on five biased-but-competent cheap judges and is a no-op on frontier ones (Spearman $ρ{=}{-}0.84$ between competence and gain over the $14$ competent judges, $p{<}10^{-3}$), concentrating the benefit where at-scale evaluation happens.