发表机构
NVIDIA(英伟达)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
Lightning OPD 2.0通过交叉拟合风格残差化缓解大型推理模型跨教师在线策略蒸馏的风格偏差,在数学推理与代码生成基准中优于原方法,实现了跨教师设置下的性能提升。
AI 中文摘要
在线策略蒸馏(OPD)从教师模型提供密集的token级监督,但其有效性取决于教师一致性,即提供OPD监督的模型应生成过用于训练监督微调(SFT)参考的演示。然而,当SFT数据来源混杂或未知,或SFT数据生成与后续蒸馏更适合选用不同模型时,实践中常违反该条件。在这类跨教师设置中,即使更强的OPD教师也可能仅带来远低于SFT参考的提升。我们发现,原始教师与参考的分歧包含潜在有用的特定上下文教师证据,以及与措辞、格式和推理节奏差异相关的重复成分。我们引入Lightning OPD 2.0,其采用交叉拟合风格残差化,通过rollout级交叉拟合将该重复成分估计为风格token偏差的操作代理,并在构建token级OPD更新前将其减去。在数学推理与代码生成基准测试中,Lightning OPD 2.0在跨教师设置中始终优于Lightning OPD。以Klear-Reasoner-8B-SFT为起点,Lightning OPD 2.0在AIME 2024上达到82.4%,在LiveCodeBench v5上达到63.0%。这些结果共同表明,Lightning OPD 2.0是一种实用的跨教师OPD方法,它将教师一致性作为前提条件放宽,允许独立选择SFT数据生成器和蒸馏教师。代码将很快发布。
英文摘要
On-policy distillation (OPD) provides dense token-level supervision from a teacher, but its effectiveness can depend on teacher consistency, meaning that the model providing OPD supervision should also have generated the demonstrations used to train the supervised fine-tuning (SFT) reference. However, this condition is frequently violated in practice when SFT data have mixed or unknown provenance or when different models are preferred for SFT data generation and subsequent distillation. In such cross-teacher settings, even a stronger OPD teacher can yield little improvement over the SFT reference. We find that raw teacher--reference disagreement contains potentially useful context-specific teacher evidence as well as a recurring component associated with differences in wording, formatting, and reasoning cadence. We introduce Lightning OPD 2.0 with cross-fitted style residualization, which uses rollout-level cross-fitting to estimate this recurring component as an operational proxy for style-token bias and subtracts it before constructing the token-level OPD update. Across mathematical reasoning and code generation benchmarks, Lightning OPD 2.0 consistently outperforms Lightning OPD in cross-teacher settings. Starting from Klear-Reasoner-8B-SFT, Lightning OPD 2.0 reaches 82.4% on AIME 2024 and 63.0% on LiveCodeBench v5. Together, these results establish Lightning OPD 2.0 as a practical approach to cross-teacher OPD, relaxing teacher consistency as a prerequisite and allowing the SFT data generator and distillation teacher to be selected independently. Code will be released soon.