arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

稳定答案,未竟推理:为何自一致性不是安全的早停信号

Stable Answers, Unfinished Reasoning: Why Self-Consensus Is Not a Safe Early-Exit Signal

Yunxiang Mo, Donghao Zhao, Hejia Geng

arXiv 2609.09989首次发表:更新:

发表机构

HKUST; University of Oxford(香港科技大学; 牛津大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究通过大规模预注册扫描证明,基于自一致性的早停规则在节省推理成本时不可靠,因一致性反映答案稳定性而非推理终止,故不能作为安全的早停信号。

AI 中文摘要

降低推理模型推理成本的一种自然方法是,反复探测单个部分轨迹以获取其当前答案,并在探测结果一致时停止——即自一致性。我们探究是否存在任何此类规则既安全又节省令牌,以及是否可以选择一次并重复使用。一项预先注册的扫描,涵盖3,520条一致性规则,在来自两个模型和三个基准的冻结轨迹上重放,未能通过预先设定的三个接受门槛中的任何一个;该前沿在留出分割集和两个未见模型上重现——而通过同一流程扫描的边界置信度控制(DEER)则通过了全部三个门槛。原因在于信号本身:一致性表明当前答案在固定探测程序下持续存在,而非推理已终止——即一致性-终止差距。基于此停止会提交非终止答案。在仍节省32%令牌的规则下,每九次停止中就有一次发生在轨迹本身后来放弃的答案上,且这些停止中的大多数截断了本会进行的修正。扩大一致性窗口并不能消除这些问题:其比例稳定在约7%,而到那时节省已降至8%。探测措辞的更改和手工标注的错误分类法表明,一致的答案往往是模型尚未确定的占位符。单独用作停止信号时,一致性失败并非因其不够严格,而是因为它反复测量了错误的对象。

英文摘要

A natural way to cut reasoning-model inference cost is to repeatedly probe a single partial trajectory for its current answer and stop once probes agree -- self-consensus. We ask whether any such rule is both safe and token-saving, and whether one can be selected once and reused. A preregistered sweep of 3,520 consensus rules, replayed on frozen trajectories from two models and three benchmarks, clears none of three acceptance gates fixed in advance; the frontier reproduces on a held-out split and on two unseen models -- while a boundary-confidence control (DEER) swept through the same pipeline clears all three. The reason lies in the signal: agreement establishes that the current answer persists under a fixed probing procedure, not that the reasoning has terminated -- a consensus-termination gap. Stopping on it commits non-terminal answers. At a rule still saving 32% of the tokens, one stop in nine fires on an answer the trajectory itself later abandons, and most of those stops cut off a correction it would otherwise have made. Widening the agreement window does not remove them: the share levels off near 7%, and by then the saving has fallen to 8%. Probe re-wording and a hand-labelled error taxonomy show the agreed answer is often a placeholder the model had not settled on. Used on its own as the stop signal, agreement fails not because it is insufficiently strict, but because it repeatedly measures the wrong object.

Comments21 pages, 9 figures, 10 tables. Yunxiang Mo and Donghao Zhao contributed equally. Code and data will be released at https://github.com/Antony-zdh/stable-answers-unfinished-reasoning

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

相关深度报道

↑