发表机构
ByteDance Inc.; University of California, San Diego(字节跳动有限公司; 加利福尼亚大学圣迭戈分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对视频扩散模型蒸馏中评论家误差累积导致的样本退化问题,提出投影分布匹配蒸馏(PDMD),通过投影去除误差分量,仅需一行代码改动即可稳定训练并提升样本质量,在多个基准上超越现有方法。
AI 中文摘要
现代视频扩散模型需要对长时空token序列进行数十次去噪评估。分布匹配蒸馏(DMD)将函数评估次数(NFE)减少到仅几次。然而,DMD样本在训练过程中可能会退化,表现出逐渐过饱和和伪影。我们将这种不稳定性追溯到评论家误差,这些误差进入连续的学生更新并随时间累积。我们引入了投影分布匹配蒸馏(PDMD)来过滤评论家误差。PDMD投影掉DMD更新中与学生-评论家端点残差平行的分量。在固定的噪声查询下,我们证明该残差是评论家端点误差的无偏估计。在高维假设下,这种投影去除了评论家误差的恒定比例,同时仅丢弃了理想DMD信号的消失比例。经验上,投影稳定了训练,并在DMD退化并产生不自然纹理的地方提高了样本质量。PDMD只需对DMD进行一行代码更改,无需额外的损失、网络、数据、模型传递或多阶段训练。使用Wan2.1,PDMD在4 NFE下实现了83.73的VBench总分,超过匹配的DMD 1.03分。在MiniMax-H3联合视频-音频生成中,PDMD实现了83.17的VideoGen-Eval视觉总分,比最强的蒸馏基线高0.41分。PDMD在比较的4-NFE模型中的全部六个音频指标上也取得了最佳性能。定性比较和用户研究在视觉质量、运动和音频质量方面都倾向于PDMD而非蒸馏基线。代码和模型可在https URL获得。
英文摘要
Modern video diffusion models require tens of denoising evaluations over long spatiotemporal token sequences. Distribution Matching Distillation (DMD) reduces the number of function evaluations (NFE) to just a few. However, DMD samples can degrade during training, exhibiting progressive oversaturation and artifacts. We trace this instability to critic errors, which enter successive student updates and accumulate over time. We introduce Projected Distribution Matching Distillation (PDMD) to filter critic errors. PDMD projects out the component of the DMD update parallel to the student-critic endpoint residual. At a fixed noisy query, we prove that this residual is an unbiased estimate of the critic's endpoint error. Under high-dimensional assumptions, this projection removes a constant fraction of critic error while discarding only a vanishing fraction of ideal DMD signal. Empirically, the projection stabilizes training and improves sample quality where DMD degrades and develops unnatural textures. PDMD requires only a one-line code change to DMD, with no extra loss, network, data, model pass, or multi-stage training. With Wan2.1, PDMD achieves a VBench total score of 83.73 at 4 NFE, surpassing matched DMD by 1.03 points. On MiniMax-H3 joint video-audio generation, PDMD achieves a VideoGen-Eval visual total score of 83.17, 0.41 points above the strongest distilled baseline. PDMD also achieves the best performance on all six audio metrics among the compared 4-NFE models. Qualitative comparisons and user studies favor PDMD over the distilled baselines in visual quality, motion, and audio quality. Code and models are available at https://pdmd2026.github.io/.