一致性高估证据:LLM 裁判共识中的错误依赖性
Agreement Overstates Evidence: Error Dependence in LLM Judge Consensus
浏览论文内容
中文总结 AI 辅助
本研究揭示LLM裁判共识因错误依赖性而高估证据强度,提出利用可信示例估计准确性并识别共享错误,以指导投票方法选择。
中文摘要 AI 辅助
LLM 裁判之间的共识常常被视为决策正确性的强有力证据。这假设裁判们独立地犯错误。然而在实践中,LLM 裁判通常以相似的方式被训练和评估,因此它们可能犯相同的错误。我们研究了这种依赖性如何影响共识的可靠性。我们发现,在开放权重和前沿 LLM 裁判中均存在显著的错误相关性。在我们主要的十个裁判组中,裁判错误之间的平均成对相关性为 0.21。因此,这十个裁判仅提供大约相当于 3.5 个独立裁判的统计信息。在我们评估的高精度前沿裁判中,包括来自不同提供商的裁判,这种依赖性甚至更强。在我们高达 28% 的比较中,忽略共享错误会导致得出一个系统显著更优的结论,而考虑这些错误则不会得出该结论。我们还发现,错误的模式也很重要。大多数裁判共享的错误和集中在较小群体中的错误对共识的影响不同,并且有利于不同的投票方法。因此,仅测量整体相关性程度是不够的。我们的结果提出了一种简单的方法:使用一小部分可信示例来估计裁判准确性并识别共享错误。在分析结果时应考虑这些共享错误,并且应在将投票方法应用于新数据之前,使用可信示例来选择投票方法。
英文摘要
Consensus among LLM judges is often taken as strong evidence that a decision is correct. This assumes that judges make their errors independently. In practice, LLM judges are often trained and evaluated in similar ways, so they can make the same mistakes. We study how this dependency affects the reliability of consensus. We find substantial error correlation across both open-weight and frontier LLM judges. In our main bank of ten judges, the average pairwise correlation between judge errors is 0.21. As a result, the ten judges only provide roughly as much statistical information as 3.5 independent judges. The dependency is even stronger among the high-accuracy frontier judges we evaluate, including judges from different providers. In up to 28% of our comparisons, ignoring shared errors leads to the conclusion that one system is significantly better, while accounting for them does not. We also find that the pattern of errors matters. Errors shared by most judges and errors concentrated among a smaller group affect consensus differently and favor different voting methods. Measuring the overall amount of correlation alone is therefore insufficient. Our results suggest a simple approach: use a small set of trusted examples to estimate judge accuracy and identify shared mistakes. These shared errors should then be considered when analyzing the results, and the voting method should be chosen using trusted examples before it is applied to new data.
发表机构
- University of Central Florida(中佛罗里达大学)
机构由 AI 辅助整理,请以论文原文为准。