发表机构
University of North Carolina at Chapel Hill; Brigham Young University; Microsoft(北卡罗来纳大学教堂山分校; 杨百翰大学; 微软)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究识别在线策略蒸馏中长度膨胀源于师生模型间终止令牌不匹配,提出将功能等效EOS令牌视为共享语义停止动作以缓解,并分析训练阶段演变,发布修正实现。
AI 中文摘要
我们研究了在线策略蒸馏(OPD)中的长度膨胀现象,即学生模型的响应可能变得过长,甚至耗尽生成预算。我们识别出基础学生模型与后训练教师模型之间的“终止令牌不匹配”是这一行为的重要来源。在Qwen3、Llama和Gemma中,即使两个模型声明的停止集合相同,它们也可能将停止概率放在不同的EOS令牌上。这种不匹配会抑制学生模型偏好的终止动作,而无法可靠地转移教师模型偏好的替代动作。我们表明,仅对齐解码停止集合是不够的,而将功能等效的EOS令牌视为共享的语义停止动作,能显著缓解所有三个模型家族中由不匹配引起的长度膨胀。为了进一步理解终止行为在训练过程中的演变,我们研究了不同K2-Horizon训练阶段的OPD。这种分阶段分析表明,终止偏好可能在训练期间发生显著变化,同时也揭示了OPD运行后期出现的明显长度膨胀,这种膨胀在终止对齐后仍然存在。这些结果共同表明,终止不匹配是OPD长度动态的一个重要但并非详尽无遗的来源。我们发布了一个包含所提出的终止处理修正的实现。
英文摘要
We study length inflation in on-policy distillation (OPD), where student responses can become excessively long and even exhaust the generation budget. We identify \emph{termination-token mismatch} between base students and post-trained teachers as an important source of this behavior. Across Qwen3, Llama, and Gemma, the two models can place their stopping probability on different EOS tokens, even when their declared stopping sets are identical. This mismatch can suppress the student's preferred termination action without reliably transferring the teacher-preferred alternative. We show that aligning the decoding stopping set alone is insufficient, while treating functionally equivalent EOS tokens as a shared semantic stopping action substantially mitigates mismatch-induced length inflation across all three model families. To further understand how termination behavior evolves over training, we study OPD across different K2-Horizon training stages. This stage-wise analysis shows that termination preferences can shift substantially during training, while also revealing a distinct length inflation late in the OPD run that persists beyond termination alignment. Together, these results identify termination mismatch as an important, but not exhaustive, source of OPD length dynamics. We release an implementation incorporating the proposed termination-handling corrections.
Comments30 pages, 12 figures, 3 tables, code available at https://github.com/UNCSciML/opd-eos