arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.16333cs.CLcs.AI

步骤级在线策略蒸馏:在线策略蒸馏与监督微调的插值方法

Step-Level On-Policy Distillation: Interpolating Between On-Policy Distillation and Supervised Fine-Tuning

  • State Key Laboratory for Novel Software Technology, Nanjing University(南京大学现代软件技术国家重点实验室)
  • XingYun Lab, HUJING Digital Media & Entertainment Group(星云实验室,沪景数字媒体娱乐集团)
  • University of Chinese Academy of Sciences(中国科学院大学)
  • School of Data Science, Fudan University(复旦大学数据科学学院)

机构由 AI 辅助整理,请以论文原文为准。

Changhui Sun, Lanbo Liu, Hang Lei, Tong Ling, Jiahang Xie, Zhiyong Zheng, Yujia Wang, Hao Liu, Feng Xiao, Lu Liu, Yanlong Du, Zifeng Cheng, Ziwei Jiang, Qing Gu

AI总结:

该研究针对令牌级在线策略蒸馏的局限,提出步骤级在线策略蒸馏SOPD,结合监督微调与在线策略蒸馏的优势,在推理和智能体任务上显著优于传统方法,为蒸馏研究提供新视角。

AI中文摘要:

在线策略蒸馏(OPD)让学生模型与教师模型在学生生成轨迹上的对数分布对齐,该方法已取得出色的经验效果,且通常能以少得多的数据超越传统离线策略蒸馏。不过,标准的令牌级OPD只能在错误的学生轨迹上提供碎片化修正,无法展开完整且正确的修复路径。受此局限启发,我们提出步骤级在线策略蒸馏(SOPD),它结合了监督微调(SFT)的长程修正能力与OPD的在线策略优势,能对完整的学生生成轨迹提供步骤级监督。我们证明,在不同的步骤长度极限下,SOPD可退化为SFT或近似OPD。与SFT相比,SOPD中的教师响应以学生轨迹为条件,因此与学生访问的状态对齐更紧密;与OPD相比,SOPD提供长程修正而非碎片化令牌级指导。在推理和智能体任务上,SOPD均显著优于传统SFT和OPD,例如在ALFWorld上,SOPD较普通OPD将平均成功率提升了13.4个百分点。我们希望这项工作能为蒸馏方法的未来研究提供新视角。

英文摘要:

On-policy distillation (OPD) aligns a student model with a teacher's logit distribution on student-generated trajectories. This approach has achieved strong empirical gains and can often surpass conventional off-policy distillation with substantially less data. However, standard token-level OPD can provide only fragmented corrections along an erroneous student trajectory and cannot unfold a complete and correct repair path. Motivated by this limitation, we propose \emph{Step-Level On-Policy Distillation} (SOPD), which combines the long-horizon correction of supervised fine-tuning (SFT) with the on-policy advantage of OPD to provide step-level supervision over complete student-generated trajectories. We show that, at different limits of step length, SOPD reduces to SFT or approximates OPD. Compared with SFT, the teacher responses in SOPD are conditioned on student trajectories and therefore align more closely with student-visited states; compared with OPD, SOPD provides longer-horizon corrections rather than fragmented token-level guidance. Across both reasoning and agent tasks, SOPD substantially outperforms conventional SFT and OPD. For example, on ALFWorld, SOPD improves the average success rate by 13.4 points over Vanilla OPD. We hope this work offers a new perspective for future research on distillation methods.

↑