arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

改进视频稀疏注意力:细粒度路由器与稀疏重基

Improving Video Sparse Attention with Fine-grained Router and Sparse Rebasing

Peiyuan Zhang, Guoqiang Wei, Yilong Zhao, Zixiang Zhang, Wei Zhou, Will Lin, Heng Zhang, Xiaonan Nie, Yan Zeng, Hao Zhang

arXiv 2609.32882首次发表:更新:

AI 中文总结

提出VSA2稀疏注意力机制,通过细粒度路由器与难到易课程训练,在视频DiT中减少一半注意力计算,实现8.9倍注意力加速和4.62倍端到端加速,同时保持或提升视频质量。

AI 中文摘要

我们提出了VSA2,一种面向视频DiT的前沿可训练稀疏注意力机制。VSA2包含多种新的架构特性和训练流程,这些特性和流程应用于DiT开发周期的所有阶段,包括预训练、强化学习和推理,以生成与全注意力对应模型质量相当或更优的DiT。在架构上,VSA2引入了一个细粒度路由器,提高了识别关键令牌的精度,并通过允许每个查询关注可变数量的键值对来支持动态计算。在训练中,我们识别出“从难到易课程”(Hard-to-Easy Curriculum),即在高度稀疏下训练并在推理时以较低稀疏度评估的模型,不仅能有效泛化,而且在运动质量上优于使用全注意力训练的模型。VSA2还具有灵活性:它可以在渐进式从低到高分辨率预训练的中途替换全注意力,对早期阶段的全注意力检查点进行重基。实验表明,与VSA相比,VSA2将注意力计算减少了一半,且损失更低。在720p视频上,与FlashAttention-3基线相比,它使注意力加速了8.9倍,端到端生成加速了4.62倍,同时实现了相当或更好的视频质量。

英文摘要

We present VSA2, a frontier trainable sparse attention for video DiTs. VSA2 includes a variety of new architectural features and training procedures that we apply across all stages of the DiT development cycle, including pretraining, RL, and inference, to produce a DiT with comparable or better quality than a full attention counterpart. Architecturally, VSA2 introduces a fine-grained router that improves the precision of identifying critical tokens and supports dynamic computation by allowing each query to attend to a variable number of key-value pairs. In training, we identify a Hard-to-Easy Curriculum, where models trained under high sparsity and later evaluated with lower sparsity during inference not only generalize effectively, but also outperform models trained with full attention in motion quality. VSA2 is also flexible: it can replace full attention during the middle of progressive low-to-high resolution pretraining, rebasing early-stage full-attention checkpoints. Experiments show that VSA2 reduces attention computation by half over VSA with lower loss. On 720p videos, it accelerates attention by 8.9x and end-to-end generation by 4.62x compared to the FlashAttention-3 baseline, while achieving comparable or better video quality.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑