发表机构
Apple; Duke University(苹果公司; 杜克大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究通过受控实验发现多教师在线策略蒸馏(MOPD)的准确率优势主要源于训练设计和超参数优化,而非算法本身;其训练成本为SFT的14.8至23.1倍,而简单权重合并可在几分钟内恢复专家能力,提示应优先调优离线策略基线。
AI 中文摘要
将多个从同一基础检查点开始训练的专家模型的能力进行组合,在前沿语言模型的后训练中已变得越来越普遍。近期的趋势表明,多教师在线策略蒸馏(MOPD)优于传统的离线策略方法。然而,尽管MOPD产生了更高的推理和环境交互成本,我们发现其报告的大部分准确率提升归因于某些训练设计选择和超参数优化,而非算法本身。我们在两个多教师设置、四个模型和十一个基准上进行了受控的自蒸馏研究,比较了离线策略方法(即监督微调(SFT)和软标签蒸馏)与混合教师前缀蒸馏及MOPD。我们发现所有四种方法实现了几乎相同的准确率。然而,MOPD使用的训练GPU小时数是SFT的14.8至23.1倍。我们还重新审视了最近发表的四项在线策略与离线策略蒸馏的比较,发现当SFT基线在拒绝采样的教师轨迹上训练并使用独立调优的超参数时,所报告的在线策略增益大幅缩减。作为无需训练的替代方案,我们还发现简单的权重合并方法可以在几分钟的CPU合并时间内恢复专家能力,尽管其准确率随着模型规模的减小和任务干扰的增加而下降。总体而言,我们的结果质疑了近期归因于MOPD的增益,并建议对更高效的离线策略基线进行仔细调优作为可行的替代方案。
英文摘要
Combining capabilities of multiple expert models trained starting from the same base checkpoint has become increasingly common in frontier language-model post-training. Recent trends suggest that multi-teacher on-policy distillation (MOPD) outperforms conventional off-policy methods. However, despite the higher inference and environment interaction costs incurred by MOPD, we find that much of its reported accuracy gain is due to certain training design choices and hyperparameter optimization, rather than the algorithm itself. We conduct a controlled self-distillation study across two multi-teacher settings, four models, and eleven benchmarks, comparing off-policy methods, namely supervised fine-tuning (SFT) and soft-label distillation, to hybrid teacher-prefix distillation and MOPD. We found that all four methods achieve \textit{nearly identical} accuracy. However, MOPD uses $14.8$--$23.1\times$ SFT's training GPU-hours. We also revisit four recently published comparisons between on-policy and off-policy distillation and find that the reported on-policy gains shrink substantially when SFT baselines are trained on rejection-sampled teacher trajectories and use independently tuned hyperparameters. As a training-free alternative, we also find that simple weight-merging methods can recover expert capabilities with minutes of CPU merging time, although their accuracy degrades as model size decreases and task interference increases. Overall, our results question recent gains reported due to MOPD and suggest careful tuning of more efficient off-policy baselines as a viable alternative.