AI 中文总结
本研究提出 Semi-OPD 作为 OPD 的替代蒸馏方法,在 17 组不同规模的教师-学生对中多数场景下表现更优,其效果取决于教师与学生的输出 token 重叠率。
AI 中文摘要
在线策略蒸馏(On-policy distillation, OPD)已成为将教师模型的能力迁移至学生模型的热门方法。本研究提出关键问题:在线策略采样是否始终有益于任意教师-学生对的蒸馏?我们表明,一种简单的替代方法 Semi-OPD(从初始学生生成的离线 rollout 中进行蒸馏)在准确率和训练效率上通常优于 OPD。在参数规模从 15 亿到 2350 亿的 17 组教师-学生对中,Semi-OPD 在 14 组中优于 OPD,准确率提升最高达 13.6%,训练速度提升最高达 11.4 倍。我们进一步发现,OPD 与 Semi-OPD 的选择取决于初始教师与学生的对齐程度,该程度由输出 token 重叠率量化:仅当二者高度对齐且重叠率高时,OPD 才有益。深入研究表明,有效的蒸馏需要同时针对学生和教师的在线策略性;对于未对齐的对,随着上下文长度增加,学生 rollout 相对于教师会越来越偏离策略,从而削弱蒸馏信号。相比之下,Semi-OPD 通常更稳定,因为它在更短的上下文上进行蒸馏,同时覆盖完整轨迹,让学生接触到更多教师偏好的 token。除了提出 Semi-OPD 作为高效替代方案,本研究还促使学界重新思考何时使用 OPD,并研究针对有意义的教师-学生对的更强 OPD 变体。
英文摘要
On-policy distillation (OPD) has become increasingly popular for transferring teacher capabilities to student models. In this work, we ask a critical research question: Is on-policy sampling always beneficial for distilling arbitrary teacher-student pairs? We show that a simple alternative, Semi-OPD, which distills from offline rollouts generated by the initial student, can often outperform OPD in both accuracy and training efficiency. Across 17 teacher-student pairs ranging from 1.5B to 235B parameters, Semi-OPD outperforms OPD in 14 cases, with up to +13.6% accuracy and 11.4x training speedup. We further find that the choice between OPD and Semi-OPD depends on the alignment between the initial teacher and student, quantified by an output-token overlap ratio: OPD is beneficial only when the two are highly aligned with high overlap ratios. Our deeper investigation suggests that effective distillation requires on-policyness w.r.t. both the student and the teacher. For misaligned pairs, student rollouts can become increasingly off-policy w.r.t. the teacher as context length grows, weakening the distillation signal. In contrast, Semi-OPD is often more stable, as it distills on shorter contexts while covering full trajectories and exposing the student to more teacher-preferred tokens. Beyond proposing Semi-OPD as an efficient alternative, our work motivates the community to rethink when to use OPD and to study stronger OPD variants with meaningful teacher-student pairs.