arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

知止:用于减少过度思考的片段级信用分配

Know When to Stop: Segment-Level Credit Assignment for Reducing Overthinking

Chia-Hsuan Lee, Sihui Dai, Mingyang Zhou, Isha Slavin, Hsuan Su, Shi-Xiong Zhang, Sambit Sahu, William Campbell

arXiv 2607.00482首次发表:更新:

发表机构

Capital One(第一资本)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出DASH方法,通过将推理轨迹中的最终答案候选与真实值比较,实现片段级信用分配,减少推理语言模型过度思考行为,在AIME25上准确率达50.8%。

AI 中文摘要

推理语言模型经常过度思考:生成扩展的行为链,如回避、放弃方法和自我矛盾,这些行为消耗令牌但不改进答案。我们表明,这些行为不仅仅是长度的结果;即使控制响应长度,错误轨迹也比正确轨迹表现出更高比例的非生产性自我反思。解决这个问题需要识别自我反思在何处有帮助或有害,但获取这些步骤级注释成本高昂。我们观察到,推理轨迹中的中间答案承诺可以提供廉价代理:通过将轨迹中的每个最终答案候选与真实值比较,我们可以确定后续反思是否具有生产性,而无需任何额外监督。基于这一见解,我们提出DASH(漂移感知优势塑造),它根据每个推理片段是导向正确还是远离正确来分配片段级信用。在竞赛级数学基准上,DASH在过度思考普遍存在的场景中实现了最高准确率(AIME25: 50.8% vs. 45.4% GRPO),同时减少了过度思考行为,并实现了比基线更有效的自我纠正。

英文摘要

Reasoning language models frequently overthink: generating extended chains of behaviors such as hedging, approach abandonment, and self contradiction that consume tokens without improving answers. We show that these behaviors are not merely a consequence of length; even when controlling for response length, incorrect traces exhibit higher rates of unproductive self-reflection than correct ones. Addressing this requires identifying where self-reflection helps vs hurts, but obtaining these step-level annotations is costly. We observe that intermediate answer commitments within reasoning traces can provide a cheap proxy: by comparing each final answer candidate in the trace to the ground truth, we can determine whether subsequent reflection is productive without any additional supervision. Building on this insight, we propose DASH (Drift Aware advantage SHaping), which assigns segment-level credit based on whether each reasoning segment leads toward or away from correctness. On competition-level math benchmarks, DASH achieves the highest accuracy where overthinking is prevalent (Average Accuracy: 59.45% vs. 58.1% Dr.GRPO vs. 56.95% GRPO) while reducing overthinking behaviors and achieving more productive self-correction than baselines.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑