发表机构
Hefei University of Technology; Tsinghua University; Zhipu AI(合肥工业大学; 清华大学; 智谱AI)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对扩散模型大转小在线策略蒸馏失效问题,提出Fixed-State KL度量揭示分布差距根源,并设计GFD-OPD方法,通过折叠引导减少误差放大,实现高效且最优的跨尺度蒸馏。
AI 中文摘要
在线策略蒸馏(OPD)在语言模型中已展现出两个重要能力:将大型教师模型压缩为较小的学生模型,以及将专家模型合并为单一模型。然而,现有的扩散模型OPD大多聚焦于后者,且教师与学生共享相同的骨干网络和规模。我们研究了从大型教师到小型学生的扩散模型OPD(即大转小),并发现标准方法在此场景下失效。为探究根本原因,我们提出了Fixed-State KL,一种有效且公平的度量方法,用于衡量OPD训练中扩散模型学生与教师之间的分布差距。我们首次阐明为何大转小OPD对扩散模型具有挑战性:较小的学生难以完美匹配较大教师的分布,而无分类器引导(classifier-free guidance)会累积并放大学生条件分支和无条件分支与教师对应分支之间的分布差异。为解决此问题,我们提出GFD-OPD,一种简单而有效的方法,在缩小学生-教师差距的同时避免CFG组合带来的误差放大。在大量实验中,GFD在训练效率和最终性能上均优于先前基线,在所有基准上取得了最先进的结果。
英文摘要
On-policy distillation (OPD) has demonstrated two important capabilities in language models: compressing large teachers into smaller students and merging expert models into a single model. Existing diffusion OPD, however, mostly focus on the latter, with teachers and students sharing the same backbone and scale. We investigate large-to-small diffusion opd from large teachers to a small student and find that the standard recipe fails. To find the underlying cause, we propose Fixed-State KL, an effective and fair way to measure the distribution gap between student and teacher during OPD training for diffusion models. We are the first to clarify why large-to-small OPD is challenging for diffusion models: a smaller student struggles to perfectly match the distribution of a larger teacher, while classifier-free guidance can accumulate and amplify the distributional discrepancies between the student's conditional and unconditional branches and those of the teacher. To solve this problem, we propose GFD-OPD, a simple yet effective method that reduces the student-teacher gap while avoiding the error amplification of the CFG composition. Across numerous experiments, GFD outperforms previous baselines in both training efficiency and final performance, achieving state-of-the-art results on all benchmarks.