发表机构
Tencent Inc.; Peking University(腾讯公司; 北京大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究视频生成中自注意力瓶颈问题,提出FVAttn方法,通过Top-$p$路由等技术及运行时负载均衡等策略,提升自适应稀疏注意力分布式执行效率,降低负载不平衡,实现注意力和推理加速,且视频质量有竞争力。
AI 中文摘要
视频扩散变换器处理长时空序列,使自注意力成为高分辨率视频生成的主要瓶颈。无需训练的稀疏注意力降低了成本,但在多 GPU 序列并行下,自适应 Top-$p$ 路由会导致每个头的工作量不均衡。由此产生的工作负载异质性将稀疏注意力变成了秩级掉队者问题。我们提出了 \method{},这是一个无需训练的稀疏注意力系统,可提高多 GPU 序列并行下自适应稀疏注意力的分布式执行效率。\method{} 使用 Top-$p$ 路由、Top-$k$ 安全下限和视频感知块组织作为稀疏路由前端,然后在运行时修复物化掩码。运行时负载均衡通过 P2P 通信迁移少量重负载头以缩短当前关键路径。松弛感知稀疏增强用额外的高价值块填充剩余的非关键秩松弛,而重叠则将调度和迁移开销隐藏在现有计算之后。在逐步提炼的 Wan2.2 I2V 上,\method{} 将平均负载不平衡从 1.34 降低到 1.08,并比 FlashAttention 实现了 4.41 倍的注意力加速,同时在具有竞争力的视频质量下实现了 2.02 至 2.11 倍的 DiT 推理加速。
英文摘要
Video Diffusion Transformers process long spatio-temporal sequences, making self-attention the main bottleneck in high-resolution video generation. Training-free sparse attention reduces this cost, but adaptive Top-$p$ routing creates uneven per-head workloads under multi-GPU sequence parallelism. The resulting workload heterogeneity turns sparse attention into a rank-level straggler problem. We present \method{}, a training-free sparse-attention system that improves the distributed execution efficiency of adaptive sparse attention under multi-GPU sequence parallelism. \method{} uses Top-$p$ routing, a Top-$k$ safety floor, and video-aware block organization as the sparse-routing frontend, then repairs the materialized mask at runtime. Runtime Load Balancing migrates a small number of heavy heads via P2P communication to shorten the current critical path. Slack-Aware Sparse Augmentation fills residual non-critical-rank slack with additional high-value blocks, while overlap hides scheduling and migration overhead behind existing computation. On step-distilled Wan2.2 I2V, \method{} reduces average load imbalance from 1.34 to 1.08 and delivers a $4.41\times$ attention speedup over FlashAttention, while achieving a $2.02$--$2.11\times$ DiT inference speedup with competitive video quality.