发表机构
Washington University in St. Louis; AWS AI Labs; Carnegie Mellon University; Georgia Institute of Technology(圣路易斯华盛顿大学; AWS AI 实验室; 卡内基梅隆大学; 佐治亚理工学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对在线策略蒸馏中教师面对学生生成前缀时续写性能下降的问题,提出SCOUT协同训练框架,通过可验证奖励的强化学习适配教师,提升蒸馏效果。
AI 中文摘要
在线策略蒸馏(OPD)最近成为一种有前景的后训练范式,其中学生在其自身策略生成的轨迹上,在密集的教师监督下学习。然而,OPD引入了一个基本的非对称性:尽管采样的轨迹对于学生而言是在线策略的,但对于教师而言却是离线策略的。教师通常被优化为从其自身策略生成的前缀继续,但在OPD期间,它必须转而监督由学生生成的前缀。经验上,我们发现随着这些前缀变长,教师的续写性能会下降。为了解决这个问题,我们提出了学生条件化的教师更新(SCOUT),一种协同训练框架,使教师适应学生生成的前缀。在标准的OPD更新之外,SCOUT周期性地使用具有可验证奖励的强化学习来优化教师的条件能力,其中教师从学生前缀生成续写并从结果奖励中学习。受控实验表明,SCOUT提高了教师从学生生成的前缀继续的能力,支持了学生条件化教师适应的预期机制。在多种教师-学生配置、模型规模和推理领域中,SCOUT也一致地提高了在线策略蒸馏的有效性。
英文摘要
On-policy distillation (OPD) has recently emerged as a promising post-training paradigm in which the student learns from trajectories generated by its own policy under dense teacher supervision. However, OPD introduces a fundamental asymmetry: although the sampled trajectories are on-policy for the student, they are off-policy for the teacher. The teacher is typically optimized to continue from prefixes generated by its own policy, but during OPD it must instead supervise prefixes generated by the student. Empirically, we find that its continuation performance degrades as these prefixes grow longer. To address this issue, we propose Student-COnditioned Updates of the Teacher (SCOUT), a co-training framework that adapts the teacher to student-generated prefixes. Alongside standard OPD updates, SCOUT periodically optimizes the teacher's conditional ability using reinforcement learning with verifiable rewards, where the teacher generates continuations from student prefixes and learns from outcome rewards. Controlled experiments show that SCOUT improves the teacher's ability to continue from student-generated prefixes, supporting the intended mechanism of student-conditioned teacher adaptation. Across multiple teacher--student configurations, model scales, and reasoning domains, SCOUT also consistently improves the effectiveness of on-policy distillation.