发表机构
KAIST AI; Kakao(韩国科学技术院人工智能学院; Kakao公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对多教师在线蒸馏中初始化影响恢复的问题,提出迭代合并方法,从均匀合并逐步添加任务向量,在5领域设置中实现更高平均归一化恢复。
AI 中文摘要
多教师在线蒸馏(MOPD)将独立开发的领域教师模型合并为一个学生模型,通过蒸馏它们在学生生成样本上的预测来实现。我们研究了一种设置,其中教师模型共享一个参考模型,但经历了不同的后训练流程,并发现MOPD可能难以恢复某些教师能力。由于蒸馏发生在学生生成的前缀上,学生初始化会强烈影响后续的恢复效果。然而,初始基准性能并不是MOPD良好初始化的可靠预测指标。例如,合并初始化可能从低于SFT预热开始,但在MOPD后最终达到更高水平。我们进一步发现,有效的合并既依赖于相对教师贡献,也依赖于整体合并规模,一些强配置位于凸参数平均单纯形之外。因此,选择良好的合并初始化不仅需要评估其即时性能,还需要评估其在MOPD下所促进的学习,这使得一次性系数搜索变得困难。我们提出了面向MOPD的迭代合并(IM-MOPD),该方法从均匀合并开始,在蒸馏过程中逐步为恢复不足的领域添加任务向量增量。在5个领域的设置中,IM-MOPD相比使用均匀合并初始化或SFT预热的MOPD,实现了更高的平均归一化恢复,表明有效的教师贡献可以在训练过程中逐步确定。
英文摘要
Multi-teacher on-policy distillation (MOPD) combines independently developed domain teachers into a single student by distilling their predictions on student-generated samples. We study a setting where teachers share a reference model but undergo different post-training procedures, and find that MOPD can struggle to recover some teacher capabilities. Because distillation occurs on student-generated prefixes, the student initialization can strongly affect subsequent recovery. However, initial benchmark performance is not a reliable predictor of a good MOPD initialization. For example, merge initialization can start below SFT warm-up yet finish higher after MOPD. We further find that effective merging depends on both the relative teacher contributions and the overall merge scale, with some strong configurations lying outside the simplex of convex parameter averaging. Thus, selecting a good merge initialization requires evaluating not only its immediate performance but also the learning it enables under MOPD, making one-shot coefficient search difficult. We propose Iterative Merging for MOPD (IM-MOPD), which starts from a uniform merge and progressively adds task-vector increments for under-recovered domains during distillation. In a 5-domain setting, IM-MOPD achieves higher average normalized recovery than MOPD with either uniform merge initialization or SFT warm-up, showing that effective teacher contributions can be determined progressively during training.