arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

更密集 ≠ 更好:持续后训练中同策略自蒸馏的局限性

Denser $\neq$ Better: Limits of On-Policy Self-Distillation for Continual Post-Training

Meng Wang, Haohan Zhao, Wenzhuo Liu, Lu Yang, Geng Liu, Haiyang Guo, Guo-Sen Xie, Gaofeng Meng, Hongbin Liu, Fei Zhu

arXiv 2607.01763首次发表:更新:

AI 中文总结

研究通过自蒸馏策略优化(SDPO)发现,同策略自蒸馏在持续后训练中会加剧遗忘甚至崩溃,而GRPO等强化学习方法更保守地保留先前能力,表明密集自蒸馏并非默认的稳定器。

AI 中文摘要

持续后训练使基础模型能够获取新知识同时保留现有能力。近期研究表明,同策略学习可以缓解遗忘,其中同策略自蒸馏成为一种特别有吸引力的方法。本文通过自蒸馏策略优化(SDPO)重新审视了这一乐观观点。我们的实验表明,当教师信号稳定且对齐良好时,SDPO可以加速领域内特化,但难以泛化到分布外场景。在持续后训练中,SDPO表现出更强的遗忘甚至可能崩溃,而GRPO等同策略强化学习方法则更保守地适应并更好地保留先前能力。进一步分析揭示,更密集的自蒸馏在参数空间和响应空间中均导致更大的漂移,并通过自我强化的师生循环放大高频格式伪影。这些发现表明,仅靠同策略数据不足以进行持续学习。当教师目标稳定且令牌级监督可靠时,密集自蒸馏可以加速特化,但不应将其视为持续后训练的默认稳定器。我们的代码可在该 https URL 获取。

英文摘要

Continual post-training enables foundation models to acquire new knowledge while preserving existing capabilities. Recent work suggests that on-policy learning can mitigate forgetting, with self-distillation as a particularly attractive approach. We revisit this optimistic claim through self-distillation policy optimization (SDPO). Our experiments show that SDPO accelerates in-domain specialization when teacher signals are stable and well aligned, but struggles to generalize out of distribution. In continual post-training, SDPO exhibits greater forgetting and can even collapse, whereas GRPO, the more established on-policy reinforcement learning method, adapts more conservatively and better preserves prior capabilities. Further analyses link these failures to increased drift in parameter and response space, and to amplification of high-frequency artifacts through a self-reinforcing teacher-student loop. Thus, on-policy data alone is insufficient for continual learning. Self-distillation is effective when teacher targets are stable and token-level supervision is reliable, but should not be treated as a default stabilizer for continual post-training. Our code is available at https://github.com/Moenupa/SDPO-CL.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑