AI 中文总结
本研究通过受控单教师OPD比较RL与SFT教师,发现RL教师指导的学生在多个任务上表现更优,且RL教师更接近共享初始化,更易被学生跟随。
AI 中文摘要
从共享检查点训练的领域专家可以通过在线策略蒸馏(OPD)合并为一个模型,其中专家作为教师,在学生自身的轨迹上监督学生。一个上游选择很少被审视:是用监督微调(SFT)还是强化学习(RL)来构建每个专家。然而,同等强大的教师未必是同等优秀的教师。我们通过受控的单教师OPD来探究这一选择,这是多教师OPD的构建模块:在Agentic、Reasoning和Perception三个领域,从Qwen3.5-9B训练出性能相当的SFT和RL教师,每个教师指导一个从该模型初始化的学生。在它们的最佳检查点上,RL指导的学生在Agentic、Reasoning和Perception上分别比SFT指导的学生高出4.27、1.50和0.86个百分点,并且恢复了更多教师相对于基础模型的性能提升。这种对比在Agentic中最为明显,其中最佳的SFT指导的学生仅恢复了其教师增益的44.44%,而最佳的RL指导的学生恢复了115.00%,超过了其教师。我们的分析指向一个解释:RL教师在参数空间中比SFT教师更接近共享初始化,因此学生更容易跟随。
英文摘要
Domain experts trained from a shared checkpoint can transfer their specialized capabilities to a single student through on-policy distillation (OPD). Existing research primarily focuses on improving this merging process, while the algorithms used to train the experts have received limited systematic comparison. We investigate which training algorithm produces teachers better suited to OPD through controlled single-teacher comparisons of supervised fine-tuning (SFT) and reinforcement learning (RL) across Agentic, Reasoning, and Perception. Teachers and students share the same Qwen3.5-9B initialization, and the two teacher types are compared at similar task performance. Our experiments show that RL teachers yield stronger students and higher recovery of teacher performance gains across all three domains. At their best checkpoints, RL-guided students outperform SFT-guided students by 4.27, 1.50, and 0.86 percentage points, respectively. In Agentic, the best SFT-guided student recovers only 44.44% of its teacher's performance gain over the base model, whereas the best RL-guided student recovers 115.00%, surpassing its teacher. Our further analysis shows that RL teachers undergo smaller parameter displacements from the shared initialization than SFT teachers. These findings support the hypothesis that RL teachers' smaller departures from the student's starting point facilitate learning through OPD, resulting in stronger students.