理解SFT、RLVR与OPD在LLM后训练中的协同作用
Understanding the Synergy between SFT, RLVR, and OPD in LLM Post-Training
浏览论文内容
中文总结 AI 辅助
本研究通过Qwen3实验揭示SFT、RLVR与OPD在后训练中的协同效应,发现阶段间兼容性决定整体效果,教师调整与学生预热可显著提升蒸馏性能。
中文摘要 AI 辅助
现代LLM后训练将监督微调(SFT)、可验证奖励强化学习(RLVR)和在线策略蒸馏(OPD)组合成多阶段流水线,然而这些阶段通常被孤立地设计和评估。我们表明这种组合具有重要影响:一个改进当前模型的阶段可能使下一阶段效果降低。通过使用Qwen3模型在数学和科学推理任务上的受控实验,我们首先在9个师生对(参数比从2倍到53倍)中刻画了OPD的特性,并表明OPD的有效性取决于师生兼容性,而非仅取决于教师规模。OPD的周围阶段以三种方式重塑这种兼容性:(1)短暂的SFT预热改善了后续OPD,而经过RLVR强化的学生在同一教师的蒸馏下表现退步。(2)用RLVR调整教师会按其所增加的能力比例提高下游OPD的准确率。在这两种干预之后,我们发现仅结合教师调整和学生预热,在相同蒸馏步数后,平均OPD准确率从29.2%提高到43.8%(相对提升50%),并伴随额外的预备训练。(3)在可比的准确率下,OPD为下游RLVR留下的初始化比SFT更强,且随着RL计算规模的扩大,这一差距也在扩大。我们的结果表明,每个后训练阶段的选择不仅应基于其增加的能力,还应基于其为下一阶段创造的学习接口。
英文摘要
Modern LLM post-training composes supervised fine-tuning (SFT), reinforcement learning with verifiable rewards (RLVR), and on-policy distillation (OPD) into multi-stage pipelines, yet these stages are typically designed and evaluated in isolation. We show that this composition is consequential: a stage that improves the current model can make the next stage less effective. Through controlled experiments with Qwen3 models on math and science reasoning, we first characterize OPD across nine student-teacher pairs spanning 2x to 53x parameter ratios and show that OPD effectiveness depends on student-teacher compatibility rather than teacher scale alone. The surrounding stages of OPD reshape this compatibility in three ways: (1) A brief SFT warm-up improves subsequent OPD, while an RLVR-strengthened student regresses under distillation from the same teacher. (2) Adapting the teacher with RLVR raises downstream OPD accuracy in proportion to the capability it adds. Following these two interventions, we find that combining teacher adaptation and student warm-up alone raise average OPD accuracy from 29.2\% to 43.8\% (50\% relative improvement) after the same number of distillation steps, with additional preparatory training. (3) At comparable accuracy, OPD leaves a stronger initialization for downstream RLVR than SFT, with a gap that widens as RL compute scales. Our results suggest that each post-training stage should be chosen not only for the capability it adds, but for the learning interface it creates for the next stage.