发表机构
Rice University; Apple(莱斯大学; 苹果公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究通过pass@k指标揭示扩散蒸馏中训练目标对分布覆盖的影响,发现分布匹配目标牺牲大预算覆盖,而一致性和轨迹目标更好地保留教师覆盖。
AI 中文摘要
扩散蒸馏被广泛用于加速采样,由此产生的少步模型被普遍认为在生成质量上可以匹配甚至超越其多步教师模型。然而,诸如GenEval2之类的标准评估通常每个提示仅抽取一个样本,因此改进的分数可能无法揭示分布覆盖的损失。因此,我们使用\textbf{pass@$\mathbf{k}$}重新审视蒸馏模型是否真正匹配其教师模型,而不仅仅是在单次抽取性能上。pass@$k$衡量在$k$个独立样本中至少有一个满足质量标准的概率。当$k{=}1$时,pass@$k$退化为标准的单次抽取评估。随着$k$的增长,该曲线揭示了额外抽取是找到真正不同的成功还是仅仅重复相同的模式,直接展示了模型覆盖有效输出空间的广度。我们首先证明,分类器自由引导(CFG)——其质量与覆盖的权衡已被充分确立——是最清晰的案例:更高的引导提高了pass@$1$,但其优势在更大的$k$下缩小并反转。将pass@$k$应用于少步蒸馏模型,我们发现同样的权衡沿训练目标分裂:分布匹配目标集中学生的输出分布,提高了早期命中率但侵蚀了大预算覆盖,而基于一致性和轨迹的目标即使在大的$k$下也能更好地保留教师的覆盖。我们进一步表明,这种权衡延伸到少步因果视频生成。我们的发现揭示了扩散蒸馏的一个先前被忽视的代价:在图像和视频生成中,训练目标的选择从根本上决定了一个少步模型是继承其教师的分布覆盖,还是为了单次抽取质量而牺牲它。
英文摘要
Diffusion distillation is widely adopted to accelerate sampling, and the resulting few-step models are broadly believed to match or even surpass their multi-step teachers in generation. However, standard evaluations such as GenEval2 typically draw only one sample per prompt, so improved scores may fail to reveal losses in distribution coverage. We therefore revisit whether distilled models truly match their teachers beyond single-draw performance using \textbf{pass@$\mathbf{k}$}, which measures the probability that at least one of $k$ independent samples satisfies a quality criterion. At $k{=}1$, pass@$k$ reduces to standard single-draw evaluation. As $k$ grows, the curve reveals whether additional draws find genuinely different successes or merely revisit the same modes, directly exposing how broadly a model covers the space of valid outputs. We first show that classifier-free guidance (CFG), whose quality--coverage tradeoff is well established, is the clearest case: higher guidance improves pass@$1$, but its advantage shrinks and reverses at larger $k$. Applying pass@$k$ to few-step distilled models, we find the same tradeoff splits along training objectives: distribution-matching objectives concentrate the student's output distribution, boosting early-hit rates while eroding large-budget coverage, whereas consistency and trajectory-based objectives better preserve the teacher's coverage even at large $k$. We further show that this tradeoff extends to few-step causal video generation. Our findings reveal a previously overlooked cost of diffusion distillation: across both image and video generation, the choice of training objective fundamentally determines whether a few-step model inherits its teacher's distribution coverage or trades it away for single-draw quality.
Commentsunder review