arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

推理未停止:虚假思维链终止的分析

</think> Doesn't Stop Reasoning: Analysis of Spurious CoT Termination

Seunghee Koh, Sungjae Choi, Minchan Kwon, Sunghyun Baek, Junmo Kim

arXiv 2609.03633首次发表:更新:

发表机构

Korea Advanced Institute of Science and Technology(韩国科学技术院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文针对大型推理模型的虚假思维链终止现象,提出用早停标记注意力偏置(EAB)增加对思维结束标记(EoT)的注意力,以减少虚假终止并缩短回答阶段长度,揭示了外部匹配思维块格式控制模型的局限性。

AI 中文摘要

思维链(CoT)推理可提升大型推理模型(LRM)在复杂任务上的表现,但常产生冗长冗余的推理轨迹。近期无训练的早停方法通过选择中间点停止推理以缩短轨迹。本文研究其中一种策略:在该点注入思维结束标记(EoT,</think>)以触发从推理到回答的过渡,发现注入的EoT并非总能引发清晰的回答阶段——在模型重新生成另一个EoT前,回答阶段的生成会继续,且该重新生成的EoT前的跨度与早停节省的推理标记数成正比,表现出持续的推理行为,本文将此现象称为虚假思维链终止,即类推理生成延续至回答阶段。本文假设对注入EoT的注意力不足是导致虚假思维链终止的原因,并用早停标记注意力偏置(EAB)验证该假设。在四个LRM、五个基准和两种早停方法的实验中,增加对注入EoT的注意力可减少虚假思维链终止并缩短回答阶段长度。这些结果揭示了通过外部匹配LRM显式思维块格式来控制LRM的局限性:插入EoT标记虽符合该格式,但本身无法保证实现预期的推理到回答过渡。本文代码可在该https网址获取。

英文摘要

Chain-of-thought (CoT) reasoning improves large reasoning models (LRMs) on complex tasks but often produces long, redundant traces. Recent training-free early-exit methods shorten these traces by choosing an intermediate point to stop reasoning. We study one such strategy that injects an end-of-think token (EoT, </think>) at this point to trigger the reasoning-to-answering transition, and find that the injected EoT does not always induce a clean answering phase. Answering-phase generation can continue before the model regenerates another EoT, with the span preceding this regenerated EoT scaling with the reasoning tokens saved by early exit and exhibiting continued reasoning behavior. We call this spurious CoT termination, where reasoning-like generation continues into the answering phase. We hypothesize that insufficient attention to the injected EoT contributes to spurious CoT termination and probe this hypothesis with Exit-token Attention Biasing (EAB). Across four LRMs, five benchmarks, and two early-exit methods, increasing attention to the injected EoT reduces spurious CoT termination and answering-phase length. These results reveal a limitation of controlling LRMs by externally matching their explicit think-block format. Inserting the EoT token conforms to this format but does not by itself guarantee the intended reasoning-to-answering transition. Our code is available at https://github.com/Seunghee-Koh/Spurious-CoT-Termination.

CommentsAccepted to EMNLP 2026 Main Conference

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑