arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.07874cs.LG

基于负策略回滚的在线策略蒸馏

On-Policy Distillation with Negative-Policy Rollouts

Jaehui Hwang, Dongyoon Han, Sangdoo Yun, Byeongho Heo

首次发表
浏览论文内容

中文总结 AI 辅助

提出 Negative-Policy OPD (NP-OPD),通过引入低性能负策略的 rollout 提供负面信号,补充教师正向监督,从而提升 on-policy distillation 在不同设置下的性能。

中文摘要 AI 辅助

On-policy distillation (OPD) 作为一种后训练方法已被广泛研究,其中学生模型在其自身的 rollout 上从更强的教师模型获得 token 级别的监督。最近的研究通过替代性的蒸馏奖励公式和教师配置改进了 OPD,而蒸馏的目标仍然集中在模仿教师上。然而,当更强的教师与学生之间的分布重叠有限时,这种正向指导可能提供不足的学习信号。在这项工作中,我们引入了 Negative-Policy OPD (NP-OPD),它通过来自性能较低、能力较弱的负策略的 rollout 来补充教师监督,该负策略作为学生的负面参考。NP-OPD 不修改蒸馏奖励公式,而是在 rollout 阶段引入负策略,持续提供负策略相对于教师更偏好的 token,以便这些 token 在整个训练过程中始终暴露于教师监督之下。这通过负策略 rollout 提供了明确的负面信号,同时保留了 OPD 中使用的正向教师监督。通过大量实验,我们表明 NP-OPD 在不同模型规模、生成模式、推理领域和不同的 OPD 变体上均改善了 OPD。此外,我们的分析表明,NP-OPD 有效抑制了负策略相对于教师更偏好的 token,并将学生从负策略中移开。这些结果支持了我们通过负策略 rollout 引入负面信号的设计,并为 OPD 中 rollout 策略的作用提供了新的见解。代码将在此 https URL 提供。

英文摘要

On-policy distillation (OPD) has been widely studied as a post-training method in which a student model obtains token-level supervision from a stronger teacher on its own rollouts. Recent studies have improved OPD through alternative distillation reward formulations and teacher configurations, while the objective of distillation remains centered on mimicking the teacher. However, when a stronger teacher has limited distributional overlap with the student, such positive guidance can provide insufficient learning signals. In this work, we introduce Negative-Policy OPD (NP-OPD), which complements teacher supervision with rollouts from a lower-performing, lower-capability negative policy that serves as a negative reference for the student. Rather than modifying the distillation reward formulation, NP-OPD introduces the negative policy at the rollout stage, continuously supplying tokens preferred by the negative policy over the teacher so that they remain exposed to teacher supervision throughout training. This provides an explicit negative signal through negative-policy rollouts while preserving the positive teacher supervision used in OPD. Through extensive experiments, we show that NP-OPD improves OPD across model scales, generation modes, reasoning domains, and different OPD variants. Furthermore, our analyses show that NP-OPD effectively suppresses tokens preferred by the negative policy over the teacher and moves the student away from the negative policy. These results support our design of introducing negative signals through negative-policy rollouts and provide new insight into the role of the rollout policy in OPD. Code will be available at https://github.com/naver-ai/np-opd.

发表机构

  • NAVER AI Lab

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑