arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

FVAttn:用于视频生成的具有运行时负载均衡的自适应稀疏注意力

FVAttn: Adaptive Sparse Attention with Runtime Load Balancing for Video Generation

Hao Liu, Chenghuan Huang, Ye Huang, Zhiying Wen, Hao Liu, Mohan Zhang, Chen Li, Ziyang Ma, Jing Lyu, Jiangsu Du

arXiv 2607.16190首次发表:更新:

发表机构

Tencent Inc.; Peking University(腾讯公司; 北京大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究视频生成中自注意力瓶颈问题,提出FVAttn方法,通过Top-$p$路由等技术及运行时负载均衡等策略,提升自适应稀疏注意力分布式执行效率,降低负载不平衡,实现注意力和推理加速,且视频质量有竞争力。

AI 中文摘要

视频扩散变换器处理长时空序列,使自注意力成为高分辨率视频生成的主要瓶颈。无需训练的稀疏注意力降低了成本,但在多 GPU 序列并行下,自适应 Top-$p$ 路由会导致每个头的工作量不均衡。由此产生的工作负载异质性将稀疏注意力变成了秩级掉队者问题。我们提出了 \method{},这是一个无需训练的稀疏注意力系统,可提高多 GPU 序列并行下自适应稀疏注意力的分布式执行效率。\method{} 使用 Top-$p$ 路由、Top-$k$ 安全下限和视频感知块组织作为稀疏路由前端,然后在运行时修复物化掩码。运行时负载均衡通过 P2P 通信迁移少量重负载头以缩短当前关键路径。松弛感知稀疏增强用额外的高价值块填充剩余的非关键秩松弛,而重叠则将调度和迁移开销隐藏在现有计算之后。在逐步提炼的 Wan2.2 I2V 上,\method{} 将平均负载不平衡从 1.34 降低到 1.08,并比 FlashAttention 实现了 4.41 倍的注意力加速,同时在具有竞争力的视频质量下实现了 2.02 至 2.11 倍的 DiT 推理加速。

英文摘要

Video Diffusion Transformers process long spatio-temporal sequences, making self-attention the main bottleneck in high-resolution video generation. Training-free sparse attention reduces this cost, but adaptive Top-$p$ routing creates uneven per-head workloads under multi-GPU sequence parallelism. The resulting workload heterogeneity turns sparse attention into a rank-level straggler problem. We present \method{}, a training-free sparse-attention system that improves the distributed execution efficiency of adaptive sparse attention under multi-GPU sequence parallelism. \method{} uses Top-$p$ routing, a Top-$k$ safety floor, and video-aware block organization as the sparse-routing frontend, then repairs the materialized mask at runtime. Runtime Load Balancing migrates a small number of heavy heads via P2P communication to shorten the current critical path. Slack-Aware Sparse Augmentation fills residual non-critical-rank slack with additional high-value blocks, while overlap hides scheduling and migration overhead behind existing computation. On step-distilled Wan2.2 I2V, \method{} reduces average load imbalance from 1.34 to 1.08 and delivers a $4.41\times$ attention speedup over FlashAttention, while achieving a $2.02$--$2.11\times$ DiT inference speedup with competitive video quality.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑