发表机构
University of British Columbia; Vector Institute for AI; Canada CIFAR AI Chair(不列颠哥伦比亚大学; 向量人工智能研究所; 加拿大CIFAR人工智能主席)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出一致性分布匹配,一种无模拟且无数据的蒸馏方法,通过统一样本生成与分数估计,仅用教师和学生两个模型优化单一目标,在ImageNet 256×256上以1-NFE达到FID 2.04,4-NFE达到1.37,超越无数据蒸馏基线。
AI 中文摘要
流模型和扩散模型因计算昂贵的数值积分而遭受推理速度慢的问题。蒸馏为学生模型从教师模型的动态中学习提供了一种有前景的方式,使得一步或几步生成成为可能。然而,现有方法往往依赖于精心策划的蒸馏数据集、昂贵的教师模型展开或辅助代理网络,这使模型训练和扩展变得复杂。在这项工作中,我们提出了一致性分布匹配,一种无模拟且无数据的蒸馏方法,用于加速扩散和流模型,同时保持强大的生成能力。我们的关键见解是将样本生成和分数估计统一到一个学生网络中。因此,我们的框架仅使用两个模型,一个冻结的教师模型和一个可训练的学生模型,并优化一个目标。我们证明,最小化我们的目标意味着学生流映射推送测度向教师边际分布的Wasserstein收敛。在ImageNet 256×256上,我们的方法在40个训练周期内,通过单次函数评估(1-NFE)达到了2.04的FID,4-NFE的FID为1.37,超越了无数据的最先进的蒸馏基线。我们的代码和模型可在该https URL获得。
英文摘要
Flow and diffusion models suffer from slow inference due to computationally expensive numerical integration. Distillation provides a promising way for a student model to learn from a teacher's dynamics, enabling one-step or few-step generation. However, existing methods often depend on curated distillation datasets, costly teacher rollouts, or auxiliary proxy networks, which complicate model training and scaling. In this work, we propose Consistent Distribution Matching, a simulation-free and data-free distillation method for accelerating diffusion and flow models while preserving strong generative capacity. Our key insight is to unify sample generation and score estimation with one student network. Thus, our framework uses only two models, a frozen teacher and a trainable student, and optimizes one objective. We prove that minimizing our objective indicates Wasserstein convergence of the student flow-map pushforwards to the teacher marginals. On ImageNet 256$\times$256, our method attains an FID of 2.04 with a single function evaluation (1-NFE) and a 4-NFE FID of 1.37 within 40 epochs of training, surpassing the state-of-the-art distillation baselines without data. Our code code and model are available at https://consistentdmd.github.io/.