arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Simple-OPD:揭开策略内蒸馏的预热阶段之谜

Simple-OPD: Demystifying Warm-up for On-policy Distillation

Tao Liu, Taiqiang Wu, Mao Zheng, Xuan Luo, Runming Yang, Xuewei Yang, Junjie Wang, Yujiu Yang

arXiv 2608.06802首次发表:更新:

发表机构

Tsinghua University; The University of Hong Kong; Tencent(清华大学; 香港大学; 腾讯)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文通过数据与训练视角研究OPD的预热阶段,发现预热依赖教师兼容CoT监督、LoRA近饱和训练优于全参数SFT,提出Simple-OPD方法并验证其有效性与鲁棒性。

AI 中文摘要

策略内蒸馏(On-policy distillation, OPD)是让学生模型基于自身的rollout结果进行训练,同时接收教师模型提供的token级监督,但其效果很大程度上取决于OPD之前的预热阶段。本文从数据和训练两个角度对OPD的预热阶段进行了研究。在数据方面,研究发现有效的预热依赖于与教师兼容的思维链(chain-of-thought, CoT)监督,甚至教师生成的错误rollout也能带来与正确rollout相当的益处,这表明预热主要是传递与教师兼容的思维模式,而非仅仅是正确答案。在训练方面,研究表明,与全参数的SFT相比,采用接近饱和训练时长的低秩适配(low-rank adaptation, LoRA)能更好地平衡领域内适配与分布外泛化。基于这些发现,本文提出了Simple-OPD,这是一种即插即用的初始化方法,在OPD之前通过LoRA对学生模型进行教师生成的CoT预热。在多种设置下开展的实验证明了Simple-OPD的有效性与鲁棒性。

英文摘要

On-policy distillation (OPD) trains a student on its own rollouts with token-level supervision from teacher models, but its effectiveness can depend strongly on the warm-up stage before OPD. In this paper, we demystify warm-up for OPD from both data and training perspectives. For data, we find that effective warm-up relies on teacher-compatible chain-of-thought supervision, and that even incorrect teacher rollouts can provide comparable benefits to correct ones. This suggests that warm-up primarily transfers a teacher-compatible thinking pattern rather than merely correct answers. For training, we show that low-rank adaptation (LoRA) with a near-saturation training duration better balances in-domain adaptation and out-of-distribution generalization than full-parameter SFT. Based on these findings, we propose Simple-OPD, a plug-and-play initialization method that warms up the student on teacher-generated CoT with LoRA before OPD. Experiments across diverse settings demonstrate the effectiveness and robustness of Simple-OPD.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑