arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

对关键投票的忽视:聚合独立性指标遗漏了验证真正有帮助的地方

Blind to the Pivotal Vote: Aggregate Independence Metrics Miss Where Verification Actually Helps

Yang Shu

arXiv 2608.06940首次发表:更新:

发表机构

Zhejiang University(浙江大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究发现 LLM 评委组的准确率提升集中在关键查询上,提出的按边际分层的单票替换规则可有效提升整体准确率,且总体依赖性诊断与分层效用互补。

AI 中文摘要

LLM 评委组是一种标准评估工具,但已有研究报告显示评委组的错误高度相关:9 名评委提供的有效信息大致相当于 2 名独立评委的信息,聚合仅能缩小一小部分差距。一种自然的补救措施——来自不同证据源的信号,例如执行测试套件——在大规模情况下对评委组的有效投票数未产生可区分的变化(-0.04,95% 置信区间 [-0.10, +0.02])。聚合依赖性和条件决策效用是不同的问题。基础多数算术确定了单票替换的受影响集合:仅得票差为 1 的决策可以改变。实证问题在于评委组错误率是否上升,且有用的替换是否集中在这些决策上。答案是肯定的:全部准确率提升集中在这些关键查询上,其提升幅度很大(在三个核心配置中为 +10.4 至 +23.3 个百分点),而在其他地方则恰好为零。我们在三个代码基准和四个评委组规模(9 名评委的扩展以及 56 次相关子抽样检查,提升幅度为 +6.5 至 +16.1 个百分点)上证实了这一模式。在 HumanEval+/MBPP+ 上,多数方替换规则在仅对 16.2% 的查询调用该信号的情况下,将整体准确率从 82.44% 提升至 85.62%;仅使用信号的方法则达到更高的 87.60%。因此,总体层面的依赖性诊断与按边际分层的效用是互补的,且受影响集合的特征为任何指定的单票替换策略提供了调用减少规则。

英文摘要

LLM judge panels are a standard evaluation tool, but prior work reports highly correlated panel errors: nine judges provide roughly the effective information of two independent ones, and aggregation closes only a small fraction of the gap. A natural remedy--a signal from a different evidence source, e.g., executing a test suite--produced no distinguishable change in the panel's effective-vote count at scale (-0.04, 95\% CI [-0.10, +0.02]). Aggregate dependence and conditional decision utility are different questions. Elementary majority arithmetic fixes the affected set for single-ballot substitution: only decisions with a one-vote margin can change. The empirical question is whether panel error rates rise and useful substitutions concentrate there. They do: the entire accuracy gain concentrates on these pivotal queries, where it is large (+10.4 to +23.3 percentage points across three headline configurations), and is exactly zero elsewhere. We confirm the pattern across three code benchmarks and four panel sizes (a 9-judge extension and 56 dependent subsampling checks, gain +6.5 to +16.1 percentage points). On HumanEval+/MBPP+, a majority-side replacement rule raises overall accuracy from 82.44\% to 85.62\% while invoking the signal on 16.2\% of queries; signal-only remains stronger at 87.60\%. Thus population-level dependence diagnostics and margin-stratified utility are complementary, and the affected-set characterization yields a call-reduction rule for any specified single-ballot substitution policy.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑