发表机构
Peking University; Melon Group; Tsinghua University; Alibaba Group; University of Electronic Science and Technology of China; Harbin Institute of Technology(北京大学; Melon Group; 清华大学; 阿里巴巴集团; 电子科技大学; 哈尔滨工业大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对视频扩散Transformer的高稀疏陷阱,提出SparkDiffusion统一加速框架,通过稀疏预热、轨迹混合蒸馏和FP8量化,实现单GPU高达265倍加速并保持高质量。
AI 中文摘要
视频扩散Transformer因注意力机制主导长时空令牌序列而计算成本高昂。我们识别出“高稀疏陷阱”:在极端注意力稀疏度下,逐步局部训练损失持续下降,而终端生成质量停滞甚至退化。该陷阱本质上是监督问题:主要的终端误差源于高噪声结构生成阶段,而终端对齐训练能够纠正逐步局部训练大幅扩展也无法纠正的终端误差。这引出一个简单的分阶段原则:先将稀疏架构适配为粗略先验,再纠正终端分布。我们将该原则实例化为SparkDiffusion,一个统一的视觉生成加速框架,结合了短时稀疏预热、少步轨迹混合蒸馏以及融合内核的FP8量化。SparkDiffusion在Wan2.1/Wan2.2骨干网络及T2V/I2V任务的长序列720P生成中维持97%的注意力稀疏度且保持强视觉质量,在Wan2.1-T2V-1.3B-480P上维持90%稀疏度。通过3步无CFG推理,SparkDiffusion在单块RTX 5090上对Wan2.1-T2V-14B-720P实现相比50步CFG密集基线265倍的端到端加速(H100上为220倍),并在1.3秒内完成Wan2.1-T2V-1.3B-480P视频的去噪。
英文摘要
Video diffusion transformers are expensive because attention dominates long spatiotemporal token sequences. We identify the \emph{high-sparsity trap}: at extreme attention sparsity, step-local training losses keep decreasing while terminal generation quality stagnates or degrades. The trap is one of supervision: the dominant terminal errors originate in the high-noise structure-generation stage, and terminal-aligned training corrects terminal errors that substantially extended step-local training cannot. This yields a simple staging principle: \emph{first adapt the sparse architecture into a coarse prior, then correct the terminal distribution}. We instantiate the principle as \method, a unified acceleration framework for visual generation that combines a short sparse warm-up, few-step trajectory-mixed distillation, and FP8 quantization with fused kernels. \method sustains $97\%$ attention sparsity with strong visual quality on long-sequence 720P generation across Wan2.1/Wan2.2 backbones and T2V/I2V tasks, and $90\%$ sparsity on Wan2.1-T2V-1.3B-480P. With 3-step CFG-free inference, \method achieves a $265\times$ end-to-end speedup over the 50-step CFG dense baseline for Wan2.1-T2V-14B-720P on a single RTX~5090 ($220\times$ on H100), and denoises a Wan2.1-T2V-1.3B-480P video in $1.3$s.
CommentsCode and weights are available at:https://github.com/AlibabaResearch/SparkDiffusion and https://huggingface.co/collections/alibabagroup/sparkdiffusion