发表机构
NAVER AI Lab(NAVER人工智能实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文重新审视后训练中的完整推理轨迹,发现其收益有限,而部分轨迹在重度截断下仍有效,并证明基于端点的训练可提升推理及强化学习等后训练方法。
AI 中文摘要
大型语言模型(LLMs)通常会在预先收集的推理轨迹上进行后训练,以提升其推理能力。由于推理路径复杂且相互交织,这些轨迹往往较长,常常包含通往答案过程中的绕路。然而,关于LLMs在后训练(如监督微调,SFT)中是否确实受益于学习完整轨迹,这一问题尚未得到充分探索。从我们的初步研究开始,我们发现完整轨迹仅带来有限的收益,而部分轨迹即使在重度截断下也依然有效。我们通过基于注意力的分析和受控的token移除研究来剖析推理轨迹中的冗余性,两者均表明中间token对最终推理质量的贡献微乎其微。这表明,在已知轨迹端点的情况下,避免冗余信息可能使LLMs能够利用其内部知识推断缺失步骤,从而在内部推理出连贯的替代方案。此外,我们表明,使用端点训练LLMs会导致推理行为的一致变化,并且这也惠及基于强化学习或在线策略蒸馏的后训练方法,凸显了重新审视完整推理轨迹的必要性。代码可在以下网址获取:此https URL。
英文摘要
Large language models (LLMs) are often post-trained on pre-collected reasoning trajectories to improve their reasoning capability. Such trajectories tend to be long due to complex, interwoven paths, which often include detours on the path toward the answer. However, it has been underexplored whether LLMs indeed benefit from learning complete trajectories in post-training, such as supervised fine-tuning (SFT). Starting from our pilot study, we find that full trajectories provide only limited benefit, while partial trajectories are effective even under heavy truncation. We analyze redundancy in reasoning trajectories through attention-based analyses and controlled token-removal studies, both of which show that intermediate tokens contribute minimally to final reasoning quality. This suggests that avoiding redundant information may allow LLMs to internally infer coherent alternatives by inferring missing steps from their internal knowledge, given known trajectory endpoints. Furthermore, we show that training LLMs using endpoints leads to consistent changes in reasoning behavior, and that it also benefits post-training methods based on reinforcement learning or on-policy distillation, highlighting the need to revisit complete reasoning traces. Code is available at https://github.com/naver-ai/revisiting-trace.
CommentsTo appear in EMNLP 2026 Findings