arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2505.18875cs.CV

稀疏视频生成2:通过语义感知排列加速视频生成

Sparse VideoGen2: Accelerate Video Generation with Sparse Attention via Semantic-Aware Permutation

  • University of California, Berkeley(加州大学伯克利分校)
  • MIT(麻省理工学院)
  • NVIDIA(英伟达)
  • Stanford University(斯坦福大学)

机构由 AI 辅助整理,请以论文原文为准。

Shuo Yang, Haocheng Xi, Yilong Zhao, Muyang Li, Jintao Zhang, Han Cai, Yujun Lin, Xiuyu Li, Chenfeng Xu, Jianfei Chen, Song Han, Kurt Keutzer, Ion Stoica

更新

AI总结:

本文提出SVG2框架,通过语义感知排列提升视频生成的准确性和效率,实现生成质量与效率的帕累托最优。

AI中文摘要:

Diffusion Transformers (DiTs) 是视频生成的关键,但因其注意力的二次复杂性导致显著延迟。通过仅计算关键token,稀疏注意力减少了计算成本并提供了一种有前途的加速方法。然而,现有方法在相同计算预算下无法达到最优生成质量,原因在于(1)关键token识别不准确:当前方法基于位置而非语义聚类,导致聚合表示不精确。(2)计算浪费过多:关键token分散在非关键token中,导致GPU上浪费计算,因为GPU优化处理连续token。本文提出SVG2,一种无需训练的框架,最大化识别准确性和最小化计算浪费,实现生成质量与效率的帕累托前沿权衡。SVG2的核心是语义感知排列,利用k-means根据语义相似性聚类和重新排列token。这种方法确保了精确的簇表示,提高识别准确性,并使关键token布局密集,实现高效计算而无需填充。此外,SVG2集成了top-p动态预算控制和定制内核实现,在HunyuanVideo和Wan 2.1上分别保持PSNR高达30和26,同时实现2.30倍和1.89倍的速度提升。我们的代码在https://github.com/svg-project/Sparse-VideoGen上开源。

英文摘要:

Diffusion Transformers (DiTs) are essential for video generation but suffer from significant latency due to the quadratic complexity of attention. By computing only critical tokens, sparse attention reduces computational costs and offers a promising acceleration approach. However, we identify that existing methods fail to approach optimal generation quality under the same computation budget for two reasons: (1) Inaccurate critical token identification: current methods cluster tokens based on position rather than semantics, leading to imprecise aggregated representations. (2) Excessive computation waste: critical tokens are scattered among non-critical ones, leading to wasted computation on GPUs, which are optimized for processing contiguous tokens. In this paper, we propose SVG2, a training-free framework that maximizes identification accuracy and minimizes computation waste, achieving a Pareto frontier trade-off between generation quality and efficiency. The core of SVG2 is semantic-aware permutation, which clusters and reorders tokens based on semantic similarity using k-means. This approach ensures both a precise cluster representation, improving identification accuracy, and a densified layout of critical tokens, enabling efficient computation without padding. Additionally, SVG2 integrates top-p dynamic budget control and customized kernel implementations, achieving up to 2.30x and 1.89x speedup while maintaining a PSNR of up to 30 and 26 on HunyuanVideo and Wan 2.1, respectively. Our code is open-sourced at \href{https://github.com/svg-project/Sparse-VideoGen}{https://github.com/svg-project/Sparse-VideoGen}.

↑