arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.37001cs.CVcs.AI

参数化条带注意力用于高效视频生成

Parameterized Stripe Attention for Efficient Video Generation

  • Alibaba Group(阿里巴巴集团)

机构由 AI 辅助整理,请以论文原文为准。

Xingyu Jia, Baole Ai, Ang Wang, Kang Zhao, Yong Li

AI总结:

针对视频扩散变换器推理延迟高的问题,提出参数化条带注意力PSA,统一稀疏模式并实现单内核高效处理,在HunyuanVideo和Wan 2.1上分别获得1.57倍和1.37倍加速。

AI中文摘要:

扩散变换器(DiTs)能够生成高质量视频,但存在显著的推理延迟,主要归因于计算代价高昂的全时空注意力。虽然稀疏注意力方法提供了潜在解决方案,但现有方法面临固有的灵活性与效率之间的两难困境:预定义掩码缺乏捕捉多样化注意力模式的灵活性,而运行时确定的掩码则引入开销并牺牲硬件效率。我们指出现有方法的关键局限在于缺乏对DiT注意力的统一结构化表征,并确立视频DiT注意力在时间和空间维度上均表现出周期性的对角条带结构。为了在单个高效内核中正式编码这些结构化模式,我们提出了PSA,一种参数化条带注意力,它形式化了观察到的条带规律,统一了多样化的注意力模式以实现高效掩码生成。这种统一表示使得单个硬件高效的CUDA内核能够处理所有稀疏模式,达到FlashAttention-3级别的模型FLOPs利用率。为了确定最优稀疏配置,我们提出了一种无需训练的离线搜索算法,该算法在指定误差容限下自动为每个注意力头最大化稀疏度。在HunyuanVideo和Wan 2.1上的实验表明,与FlashAttention-3基线相比,PSA实现了1.57倍和1.37倍的端到端加速,且视觉质量下降可接受。

英文摘要:

Diffusion Transformers (DiTs) enable high-quality video generation but suffer from substantial inference latency, primarily attributable to the computationally expensive full spatio-temporal attention. While sparse attention methods offer potential solutions, existing approaches face an inherent flexibility--efficiency dilemma: predefined masks lack the flexibility to capture diverse attention patterns, while runtime-determined masks introduce overheads and sacrifice hardware efficiency. We identify the lack of a unified structural characterization of DiT attention as a key limitation of existing methods, and establish that video DiT attention exhibits \textbf{periodic diagonal stripe structures} along both temporal and spatial dimensions. To formally encode these structured patterns within a single efficient kernel, we present {\bf PSA}, a parameterized stripe attention that formalizes the observed stripe regularity, unifying diverse attention patterns for efficient mask generation. This unified representation enables a single hardware-efficient CUDA kernel to process all sparse patterns, achieving FlashAttention-3-level Model FLOPs Utilization. To determine optimal sparsity configurations, we propose a training-free offline search algorithm that automatically maximizes sparsity under a specified error tolerance for each attention head. Experiments on HunyuanVideo and Wan~2.1 demonstrate that PSA achieves 1.57$\times$ and 1.37$\times$ end-to-end speedups over FlashAttention-3 baselines, with acceptable visual quality degradation.

↑