发表机构
Institute of Information Engineering, Chinese Academy of Sciences; School of Cyber Security, University of Chinese Academy of Sciences(中国科学院信息工程研究所; 中国科学院大学网络空间安全学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
CompassOPD通过去除跨族偏移并传递族内似然偏移,解决了跨族在线策略蒸馏失效问题,在多个学生族上平均推理准确率提升最高达5.50点。
AI 中文摘要
在线策略蒸馏(OPD)在教师和学生属于同一模型族时,能够对学生生成的轨迹提供密集的令牌级监督。然而,我们发现即使在分词器对齐之后,OPD在跨族设置中的有效性也会下降,且明显更强的外部教师带来的额外改进甚微。为了理解这一脱节现象,我们将跨族OPD信号分解为两个组成部分:低能力教师族参考与学生之间的偏移,以及从该参考到强教师的族内对数似然偏移。标准OPD同时传递这两个组成部分,导致偏移主导更新方向,掩盖了与教师能力提升相关的变化。我们提出CompassOPD,它去除该偏移并传递族内偏移,同时冻结的学生参考将更新锚定到学生的初始策略。因此,教师侧和学生侧的变化都在各自模型族内进行度量。在三个学生族和多个教师族上的实验表明,CompassOPD始终优于标准跨族OPD,平均推理准确率提升最高达5.50个百分点。对于MoE教师,我们通过减少专家激活直接从教师检查点构建参考,无需单独的参考检查点,同时相对于OPD保持3.43个百分点的增益。
英文摘要
On-policy distillation (OPD) provides dense token-level supervision on student-generated trajectories. Although OPD performs strongly when teacher and student belong to the same model family, we find that its effectiveness degrades in cross-family settings even after tokenizer alignment, with substantially stronger external teachers offering little additional improvement. To understand this disconnect, we decompose the cross-family OPD signal into two components: an offset between a low-capability teacher-family reference and the student, and the within-family log-likelihood shift from that reference to the strong teacher. Standard OPD transfers both components together, allowing the offset to dominate the update direction and obscure the changes associated with teacher capability improvements. We propose CompassOPD, which removes this offset and transfers the within-family shift, while a frozen student reference anchors updates to the student's initial policy. Thus, both teacher-side and student-side changes are measured within their respective model families. Experiments across three student families and multiple teacher families show that CompassOPD consistently outperforms standard cross-family OPD, improving average reasoning accuracy by up to 5.50 points. For an MoE teacher, we further construct the reference directly from the teacher checkpoint by reducing expert activation, eliminating the need for a separate reference checkpoint while retaining a 3.43-point gain over OPD.
CommentsWork in progress