arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

作为陪审团的语言模型:跨模型共识在语言模型推理方面可超越过程奖励模型

LLMs as a Jury: Cross-Model Consensus Can Outperform Process Reward Models for LLM Reasoning

Ning Liu

arXiv 2607.10139首次发表:更新:

AI 中文总结

研究在语言模型推理中从候选推理链选正确答案的问题,提出跨模型共识方法,将其视为语言模型陪审团,通过错误去相关机制工作,在多个基准测试中表现出色,能预先表征,有定律和下限。

AI 中文摘要

从候选推理链池中选择正确答案是测试时扩展的核心,但标准选择器都有成本:自一致性继承了它重新采样的单个模型的错误,而训练的奖励模型需要标记数据且分布外迁移性差。我们研究了一种推理时免费的第三信号:跨模型共识,即独立训练的模型各自解决一次问题后在最终答案上的一致程度。我们将其视为语言模型陪审团,其中一致性结构而非任何模型对其他模型的评分是验证信号。在七个基准测试中,它比自一致性能更好地选择正确答案,比模型对自身候选答案评分要好得多:在竞赛数学中它几乎能弥合与神谕选择器的差距,而自评分几乎无法缩小差距。其机制是错误去相关:独立训练的模型错误不同,所以错误答案分散而正确答案积累一致性。我们用一个无参数定律精确阐述,该定律以封闭形式推导得出,能从三个测量的陪审团统计数据预测共识准确率,平均绝对误差为0.03,并揭示了该方法的上限:存在共享错误下限,在数学上接近零,但在科学上并非微不足道。与四种训练的验证器(包括判别式、结果式和生成式奖励模型)相比,免费的语言模型陪审团在数学训练领域内与最强的验证器匹配,在领域外是最佳选择器。因此,跨模型共识是一种我们可以预先表征的验证器:一条说明何时信任它的定律,以及一个标记其无法发挥作用之处的下限。

英文摘要

Selecting the correct answer from a pool of candidate reasoning chains is the engine of test-time scaling, yet the standard selectors each carry a cost: self-consistency inherits the errors of the single model it resamples, and trained reward models need labeled data and transfer poorly off-distribution. We study a third signal, free at inference time: cross-model consensus, the degree to which independently trained models, each solving the problem once, agree on a final answer. We treat the panel as an LLM-jury, in which the verification signal is the structure of agreement itself, with no model scoring another's work. Across seven benchmarks it selects correct answers better than self-consistency and far better than a model scoring its own candidates: on competition math it closes the entire gap to an oracle selector, while self-scoring closes almost none. The mechanism is error decorrelation: independently trained models err differently, so their wrong answers scatter while the correct one accumulates agreement. We make this precise with a parameter-free law, derived in closed form, that predicts consensus accuracy from three measured panel statistics to a mean absolute error of $0.03$ and exposes the method's ceiling: a shared-error floor where models share a misconception, near zero on math but non-trivial on science. Against four trained verifiers spanning discriminative, outcome, and generative reward models, the free LLM-jury matches the strongest inside their math training domain and is the top selector outside it. Cross-model consensus is thus a verifier we can characterize in advance: a law that says when to trust it, and a floor that marks where it cannot.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑