arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2502.01776cs.CVcs.LG

Sparse VideoGen:利用时空稀疏性加速视频扩散Transformer

Sparse VideoGen: Accelerating Video Diffusion Transformers with Spatial-Temporal Sparsity

  • University of California, Berkeley(加利福尼亚大学伯克利分校)
  • MIT(麻省理工学院)
  • NVIDIA(英伟达公司)
  • Tsinghua University(清华大学)

机构由 AI 辅助整理,请以论文原文为准。

Haocheng Xi, Shuo Yang, Yilong Zhao, Chenfeng Xu, Muyang Li, Xiuyu Li, Yujun Lin, Han Cai, Jintao Zhang, Dacheng Li, Jianfei Chen, Ion Stoica, Kurt Keutzer, Song Han

更新

AI总结:

针对视频扩散Transformer计算成本高的问题,提出免训练框架Sparse VideoGen,利用3D全注意力中的时空稀疏性并结合定制内核,在保持生成质量的同时实现显著端到端加速。

AI中文摘要:

Diffusion Transformers (DiTs) 主导着视频生成领域,但其高昂的计算成本严重限制了其在现实世界中的适用性,通常即使在高性能 GPU 上也需要数十分钟才能生成几秒钟的视频。这种低效主要源于 3D Full Attention 相对于上下文长度的二次计算复杂度。在本文中,我们提出了一个名为 Sparse VideoGen (SVG) 的免训练框架,该框架利用 3D Full Attention 中固有的稀疏性来提升推理效率。我们揭示,根据不同的稀疏模式,注意力头可以动态地分为两组:(1)Spatial Head,其中每帧内仅空间相关的 token 主导注意力输出;(2)Temporal Head,其中跨不同帧的仅时间相关的 token 主导输出。基于这一洞察,SVG 提出了一种在线分析策略来捕获动态稀疏模式并预测注意力头的类型。结合新颖的硬件高效张量布局转换和定制化内核实现,SVG 在 CogVideoX-v1.5 和 HunyuanVideo 上分别实现了高达 2.28 倍和 2.33 倍的端到端加速,同时保持了生成质量。我们的代码已开源,可在 https://github.com/svg-project/Sparse-VideoGen 获取。

英文摘要:

Diffusion Transformers (DiTs) dominate video generation but their high computational cost severely limits real-world applicability, usually requiring tens of minutes to generate a few seconds of video even on high-performance GPUs. This inefficiency primarily arises from the quadratic computational complexity of 3D Full Attention with respect to the context length. In this paper, we propose a training-free framework termed Sparse VideoGen (SVG) that leverages the inherent sparsity in 3D Full Attention to boost inference efficiency. We reveal that the attention heads can be dynamically classified into two groups depending on distinct sparse patterns: (1) Spatial Head, where only spatially-related tokens within each frame dominate the attention output, and (2) Temporal Head, where only temporally-related tokens across different frames dominate. Based on this insight, SVG proposes an online profiling strategy to capture the dynamic sparse patterns and predicts the type of attention head. Combined with a novel hardware-efficient tensor layout transformation and customized kernel implementations, SVG achieves up to 2.28x and 2.33x end-to-end speedup on CogVideoX-v1.5 and HunyuanVideo, respectively, while preserving generation quality. Our code is open-sourced and is available at https://github.com/svg-project/Sparse-VideoGen

补充信息

↑