arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

面向停顿标记微调动力学的理解:模式保留视角

Towards Understanding Pause Token Fine-Tuning Dynamics: A Mode Retention Perspective

Jaehyeon Kim, Suhwan Kim, Nakyung Lee, Yeongoon Kim, Jimin Seo, Giho Lee, Jungwoo Lee

arXiv 2609.04489首次发表:更新:

发表机构

HodooAI Lab; Seoul National University(HodooAI实验室; 首尔大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究从模式保留视角探究停顿标记微调动力学,提出MBP训练规则,在Qwen和Llama模型上提升推理能力并保留通用语言理解,还可扩展至GRPO。

AI 中文摘要

停顿标记方法通过在序列中插入特殊标记来提升大语言模型(LLM)的推理能力,现有研究多从计算表达性角度解释该增益,但对停顿标记的训练动力学研究较少。本文探究停顿标记如何重塑微调的训练动力学,两项受控预实验揭示了明显的不对称性:在合成持续学习任务中,掩码停顿标记在匹配的最终适配阶段,覆盖先前学习分布的程度约低4倍(H1,模式保留);在合成数学推理探针中,边界相邻标记会编码更多下游步骤信息(H2,非近视压缩)。本文提出与两者一致的训练规则——掩码边界停顿(MBP),即放置在推理步骤边界且损失被掩码的停顿标记。在1B至8B规模的Qwen和Llama模型上,MBP持续提升推理能力,在数学任务上最高提升6个点,代码任务上最高提升2.5个点,同时保留通用语言理解能力,还证明该模式保留策略可将增益扩展至GRPO。这些结果将停顿标记重新定义为针对保留-适配权衡的训练动力学干预手段,而非仅推理时的计算设备。

英文摘要

Pause-token methods improve LLM reasoning by inserting special tokens into sequences. Prior work explains these gains through computational expressivity. However, there is relatively little investigation into the training dynamics of pause tokens. We explore how pause tokens reshape the training dynamics of fine-tuning. Two controlled pilots expose distinct asymmetries. On a synthetic continual-learning task, masked pauses overwrite a previously-learned distribution roughly 4x less at matched final adaptation (H1, mode retention); on a synthetic math-reasoning probe, the boundary-adjacent token comes to encode substantially more downstream-step information (H2, non-myopic compression). We formalize a training rule consistent with both - Masked Boundary Pause (MBP), pause tokens placed at reasoning-step boundaries with their loss masked. Across 1B-8B Qwen and Llama models, MBP consistently improves reasoning, achieving gains of up to 6 points on math and 2.5 points on code, while preserving general language understanding abilities. We further demonstrate that this mode-preserving strategy extend gains to GRPO. These results recast pause tokens as a training-dynamics intervention on the retention-adaptation trade-off, rather than merely an inference-time computation device.

Comments24 pages, 4 figures, 19 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑