arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SPADE:用于快速视频扩散模型推理的输入自适应稀疏注意力引擎

SPADE: An Input-Adaptive Sparse Attention Engine for Fast Video Diffusion Models Inference

Shanghao Liu, Renze Chen, Size Zheng, Yuanqiang Liu, Yun, Liang, Hailong Yang

arXiv 2608.03335首次发表:更新:

AI 中文总结

针对视频扩散Transformer推理开销过高的问题,提出无训练的SPADE稀疏注意力引擎,通过三部分设计实现注意力2.26x-3.40x、端到端推理1.49x-1.80x的加速且保持生成质量。

AI 中文摘要

视频扩散Transformer(vDiTs)可生成高质量内容,但存在二次自注意力开销,在视频令牌规模下推理难以实现,挑战在于输入自适应稀疏性:以可忽略的开销选择关键Q/K/V令牌并执行,以获得端到端增益。本文提出SPADE,一种无训练的稀疏注意力引擎,包含三部分:(i)vDiT-SSR,定义3D候选块并通过摘要器/估计器表达式形式化动态掩码;(ii)使用SICS和按头策略生成运行时方案;(iii)具备低开销索引搜索、闪存块稀疏注意力和内核分组的执行器。在Hunyuan-Video和Wan 2.1/2.2的文本到视频、图像到视频生成任务中,SPADE在保持质量的同时提高了稀疏度和速度,使注意力加速2.26倍至3.40倍,端到端推理加速1.49倍至1.80倍,代码已开源。

英文摘要

Video diffusion transformers (vDiTs) generate high quality but pay quadratic self-attention cost, making inference prohibitive at video-token scales. The challenge is input-adaptive sparsity: selecting critical Q/K/V tokens with negligible overhead and executing them for end-to-end gains. We present SPADE, a training-free sparse-attention engine of three parts: (i) vDiT-SSR, a specification defining 3D blocking candidates and formalizing dynamic masks via Summarizer/Estimator expressions; (ii) runtime scheme generation using SICS and a head-wise policy; and (iii) an executor with low-overhead index search, flash block-sparse attention, and kernel grouping. Across Hunyuan-Video and Wan 2.1/2.2 for text-to-video and image-to-video generation, SPADE raises sparsity and speed while preserving quality, accelerating attention by 2.26x-3.40x and end-to-end inference by 1.49x-1.80x. Our code is open-sourced at https://github.com/6somehow/DAC-SPADE.

CommentsPublished in the 63rd ACM/IEEE Design Automation Conference (DAC '26). 7 pages, 6 figures, 3 tables

Journal ref63rd ACM/IEEE Design Automation Conference (DAC '26), 2026, 7 pages

DOI:10.1145/3770743.3804015

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑