arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

当自我一致性适得其反:对于小型大型语言模型,多数投票会损害大多数硬科学问题

When Self-Consistency Backfires: Majority Vote Hurts the Majority of Hard Science Problems for Small LLMs

Utkarsh Bahuguna

arXiv 2608.11403首次发表:更新:

发表机构

Scaler School of Technology(斯凯勒科技学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究发现,多数投票的自我一致性方法会损害小型指令微调模型在多数硬科学问题上的准确率,其核心机制是置信度与正确性不匹配,且无验证器门控方法可提升该问题下的准确率。

AI 中文摘要

通过多数投票实现的自我一致性(Self-Consistency, SC)是一种广泛使用的推理时计算利用方式:采样N条思维链,返回得票最多的答案。在完整的GPQA Diamond基准测试(共198道研究生水平的科学问题)中,多数投票会降低两个不同家族的指令微调模型在大多数问题上的单问题准确率:Qwen2.5-7B模型有56.6%的问题准确率下降,Llama-3-8B模型有65.7%的问题准确率下降,其中Qwen作为主要演示对象,Llama则从接近随机的基线验证了该趋势。该效应在观察到47个探索性问题后,在151个问题的验证集上进行了预注册,且所有四个验证性假设均通过。一个网格神谕(grid oracle,将每个问题路由到{1,2,4,8,16,32,64}中最佳N值)标记了理论上的准确率上限,Qwen比N=1时高14个百分点,Llama高17个百分点,该神谕界限需要真实标签而非可部署方法。无验证器的门控方法能达到该上限:无论是多数一致门控还是令牌熵门控,在N=64的固定预算投票下,准确率提升均不超过0.002。机制很直接:在这些问题上,置信度与正确性不匹配。在最高一致度的分组中,Qwen的多数答案正确率约为一半,而Llama的该分组比其最低一致度分组准确率更低。我们预注册并在小型指令微调模型上验证了这些发现;未测试推理原生模型,我们将其标记为核心开放问题。

英文摘要

Self-consistency via majority vote reduces per-problem accuracy on most GPQA Diamond problems for small instruction-tuned models: 56.6% of problems for Qwen2.5-7B and 65.7% for Llama-3-8B. The obvious remedy is a verifier-free confidence gate. This version reports that the most natural repair also fails, and separates three signal failures that v1 treated as one. A token-entropy gate fails for a measurement reason: averaged over a chain of some 602 tokens, the statistic is a measurement of the prose rather than of confidence in the answer. On Qwen2.5-7B-Instruct-Turbo, 198 problems at 64 samples each, a sample whose answer contradicts its own problem's plurality still emits that answer at a median margin of 20.52 nats, with 75.7% above 10 nats. Both quantities were pre-registered and tested once on 69 problems no exploratory analysis had read; both passed. The unit is the whole result: pooled across the benchmark the margin separates correct from incorrect samples by +0.0604 on the fraction above 10 nats [+0.0183, +0.1017], excluding zero; per-problem and paired it does not, at -0.0168 [-0.0527, +0.0182], crossing zero. The claim is not that token log-probabilities carry no information, but that a signal with real across-question discrimination is close to useless for the within-question decision a router faces. The plurality-agreement gate's failure remains without a mechanism, and we report it as an open problem. These new claims rest on one model: a registered second-model replication was sampled and could not be evaluated, and we report that rejection rather than the result. We separately report that on hosted serverless inference at a small budget, three reasoning-native models could not be evaluated, for three separately measured reasons; all three are downloadable, so this bounds what a metered per-token API buys rather than what is knowable.

Comments19 pages, 5 figures, 4 tables. v1 accepted at the COLM 2026 Workshop on Efficient Reasoning; v2 additions are not peer reviewed. v2 revises rather than extends: Section 4.4's mechanism claim is replaced and one Discussion sentence withdrawn. All v1 results, tables and pre-registered verdicts are unchanged. Adds the answer-token margin result, a serverless reasoning wall, and ten disclosures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑