arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Settle:学习何时停止推理

Settle: Learning When to Stop Reasoning

Ryan Brown, Zihao Fu, Chris Russell

arXiv 2609.38997首次发表:更新:

发表机构

Oxford Internet Institute, University of Oxford(牛津大学牛津互联网研究所)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

Settle 通过答案稳定性学习何时停止推理,在保持准确率的同时大幅减少 token 数,并扩展了准确率与 token 数量的帕累托前沿。

AI 中文摘要

推理模型常常在答案已经稳定后仍继续生成。Settle 从已完成轨迹中的答案稳定性学习何时停止。它训练现有的推理结束标记,同时保持其他预测接近基础模型,并且推理时仅需普通解码。在 MATH-500 上使用 Qwen3-4B,Settle 将 token 数量减少 40%,准确率仅下降 0.5 个百分点。与在相同轨迹上于首个稳定答案处截断的监督微调相比,Settle 在几乎相同的 token 数量下提升了 6.16 个百分点。其停止分数能预测正确答案是否会保持正确。Settle 扩展了所评估停止方法的准确率-token 数量帕累托前沿。

英文摘要

Reasoning models often continue generating after their answers have settled. Settle learns when to stop from answer stability in completed traces. It trains the existing end-of-reasoning token while keeping other predictions close to the base model, and requires only ordinary decoding at inference. On MATH-500 with Qwen3-4B, Settle reduces token count by 40% with a 0.5-percentage-point decrease in accuracy. It gains 6.16 percentage points over supervised fine-tuning on the same traces shortened at their first stable answer, at nearly identical token counts. Its stopping score predicts whether a correct answer will remain correct. Settle extends the accuracy-token-count Pareto frontier of the evaluated stopping methods.

Comments30 pages, 4 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑