arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

WaveAlign:面向长视频生成中稀疏注意力的缓存感知查询行调度

WaveAlign: Cache-Aware Query-Row Scheduling for Sparse Attention in Long-Video Generation

Zijian Dai, Sen Han, Youhui Bai, Shannon Wang, Kan Wu, Jingkai Huang, Yuhang Wang, Jing Li, Cheng Li

arXiv 2609.34814首次发表:更新:

发表机构

University of Science and Technology of China; Institute of Artificial Intelligence, Hefei Comprehensive National Science Center; South China University of Technology(中国科学技术大学; 合肥综合性国家科学中心人工智能研究院; 华南理工大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对长视频生成中动态稀疏注意力因不规则查询行执行导致缓存效率低、加速有限的问题,提出轻量级缓存感知查询行重排序框架WaveAlign,通过两阶段优化与自适应跳过机制,在不损失生成质量的前提下显著提升缓存命中率并降低内存流量,实现推理加速。

AI 中文摘要

基于扩散Transformer(DiTs)的长视频生成会产生极长的token序列,使得注意力成为推理阶段的主要瓶颈。动态稀疏注意力可减少计算量,但实际加速效果仍有限,原因是不规则的查询行执行会降低L2缓存局部性并增加高带宽内存(HBM)流量。本文提出WaveAlign,这是一种面向动态稀疏注意力的轻量级、缓存感知的查询行重排序框架。WaveAlign将行排序建模为优化问题,并通过两个阶段进行近似求解。第一阶段推导稀疏掩码行的低秩奇异值分解(SVD)表示,将具有相似键/值(K/V)访问模式的查询行分组,以提升并发调度行之间的K/V重叠度。第二阶段利用流式GPU调度机制,按K/V块数量降序对每个波内的行排序,使当前波的短行之后紧跟下一波的长行。这一设计对齐了波边界处的K/V访问,使得共享块能在被逐出前得到复用。框架还包含一个自适应跳过模块,可避免无收益的重排序操作。由于仅对查询行和掩码行进行重排,WaveAlign保留了稀疏注意力的语义,且无需修改现有方法或后端内核。在两种GPU架构、两个视频DiT模型以及四种稀疏注意力方法上的测试表明,WaveAlign将L2缓存命中率从28.48%–36.35%提升至79.38%–89.06%,HBM读取流量最多降低92.11%,在无质量损失的前提下,内核加速比最高达1.25倍,端到端生成加速比最高达1.17倍。

英文摘要

Long-video generation with diffusion transformers (DiTs) produces extremely long token sequences, making attention a dominant inference bottleneck. Dynamic sparse attention reduces computation, but its realized speedup remains limited because irregular query-row execution degrades L2 cache locality and increases HBM traffic. We present WaveAlign, a lightweight, cache-aware query-row reordering framework for dynamic sparse attention. WaveAlign formulates row ordering as an optimization problem and approximates it with two stages. The first stage derives a low-rank SVD representation of sparse-mask rows and groups query rows with similar K/V access patterns, increasing K/V overlap among concurrently scheduled rows. The second stage exploits streaming GPU scheduling by sorting rows within each wave in descending order of their K/V-block counts, so that short rows from the current wave are followed by long rows from the next. This aligns K/V accesses across wave boundaries and enables shared blocks to be reused before eviction. An adaptive skip module avoids unprofitable reordering. By only permuting query and mask rows, WaveAlign preserves sparse-attention semantics and requires no changes to existing methods or backend kernels. Across two GPU architectures, two video DiTs, and four sparse-attention methods, WaveAlign raises the L2 cache hit ratio from 28.48%--36.35% to 79.38%--89.06%, reduces HBM read traffic by up to 92.11%, and achieves up to 1.25x kernel and 1.17x end-to-end generation speedup without quality loss.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑