发表机构
NVIDIA; Weizmann Institute of Science(英伟达; 魏茨曼科学研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对视频扩散或流模型生成计算昂贵问题,提出并行解码蒸馏方法PDD,兼容预训练模型,通过预测多去噪步骤加速生成,在多个数据集上达SOTA性能,提升视频多样性。
AI 中文摘要
视频扩散或流模型中的生成计算成本高昂,现有加速方法依赖变分分数蒸馏和对抗损失,难以优化且易出现模式崩溃。本文提出并行解码蒸馏(PDD),一种基于轨迹的简化可扩展蒸馏方法,兼容任何预训练模型,支持不同函数评估次数采样。通过预测每个网络评估的多个去噪步骤加速生成,在多个数据集上实现SOTA性能并显著提升视频多样性。
英文摘要
Generation in video diffusion or flow models is computationally expensive due to the slow and iterative sampling process. Current state-of-the-art (SOTA) acceleration methods heavily rely on variational score distillation (VSD) and adversarial losses to distill diffusion models into few-step generators. Albeit achieving high-quality video generation, these training losses are notoriously hard to optimize and suffer from mode collapse, leading to loss of video diversity and lack of motion. In this paper, we introduce Parallel Decoding Distillation (PDD), a simplified and scalable trajectory-based distillation method for fast inference of diffusion and flow matching models. Our architecture and training procedure are compatible with any pre-trained model and support sampling with a varying number of function evaluations (NFE). PDD accelerates generation by predicting multiple denoising steps per network evaluation. Conceptually, it learns a representation of the mean velocity without regressing its derivative using JVPs or finite-difference approximations. Our method achieves SOTA performance with 4-8 NFE on LTX-2.3 Text-to-Video/Audio, Wan 14B Text-to-Video, and Qwen-Image Text-to-Image. Moreover, PDD presents a significant improvement in generated video diversity.