arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

DMAD:分布匹配作为对抗蒸馏用于快速视觉生成

DMAD: Distribution Matching as Adversarial Distillation for Fast Visual Generation

Zhengming Yu, Junkun Yuan, Haotian Yang, Gordon Guocheng Qian, Yizhi Wang, Angtian Wang, Yiding Yang, Bo Liu, Xin Li, Wenping Wang, Chongyang Ma

arXiv 2610.02188首次发表:更新:

发表机构

Texas A\&M University

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出DMAD,将分布匹配蒸馏重构为对抗分类,直接学习对数密度比,无需辅助分数拟合,在图像、文本到图像和视频生成上达到领先的少步性能。

AI 中文摘要

分布匹配蒸馏(DMD)通过分别估计的目标分数与学生分数之间的差异来训练少步学生模型,因此必须额外占用内存和计算成本来维护一个辅助扩散模型,以拟合学生不断变化的分布。我们提出DMAD,即分布匹配作为对抗蒸馏,将分布匹配重新表述为分类问题,并直接学习所需的对数密度比。共享主干上的两个判别器头区分真实数据和教师样本与学生样本,对其logits的线性损失训练学生模型,无需辅助分数拟合。我们证明,在判别器最优时,这些损失通过将判别器logits与对数密度比联系起来的经典恒等式,恢复了DMD基础的分布匹配梯度。我们进一步引入基于间隔的重新加权,该机制根据真实数据头在真实样本和教师样本之间的经验logit间隔,跨噪声水平调整教师监督。DMAD在ImageNet-64x64上一步生成达到Fréchet Inception Distance(FID)1.04,在COCO-10K上四步SDXL达到14.47,在四步Wan2.1-T2V-14B上VBench总分为85.15,这些值在所比较的少步方法和多步教师中均为最佳。在MiniMax-H3-33B上,我们的四步学生在联合音视频生成中,相对于DMD2的总体人类偏好率为79.1%,相对于rCM为84.6%(不含平局)。我们的代码、模型和演示可在该https URL获取。

英文摘要

Distribution Matching Distillation (DMD) trains a few-step student from the difference between separately estimated target and student scores, so it must keep an auxiliary diffusion model fitted to the student's evolving distribution at extra memory and computation cost. We introduce DMAD, Distribution Matching as Adversarial Distillation, which recasts distribution matching as classification and learns the required log-density ratios directly. Two discriminator heads on a shared backbone distinguish real data and teacher samples from the student's, and linear losses on their logits train the student without auxiliary score fitting. We prove that at the discriminator optimum these losses recover the distribution-matching gradient underlying DMD, through the classical identity linking discriminator logits to log-density ratios. We further introduce gap-based reweighting, which adapts teacher supervision across noise levels from the real-data head's empirical logit gap between real and teacher samples. DMAD reaches a Fréchet Inception Distance (FID) of 1.04 with one-step generation on ImageNet-64x64, 14.47 with four-step SDXL on COCO-10K, and a VBench total score of 85.15 with four-step Wan2.1-T2V-14B, the best values among the compared few-step methods and the multi-step teachers. On MiniMax-H3-33B, our four-step student achieves overall human preference rates of 79.1% over DMD2 and 84.6% over rCM for joint audio-video generation, excluding ties. Our code, models and demos are available at https://yzmblog.github.io/projects/DMAD.

Comments28 pages, 15 figures. Project page: https://yzmblog.github.io/projects/DMAD

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑