解耦逻辑掩码与GPU执行以实现动态块稀疏注意力
Decoupling Logical Masks from GPU Execution for Dynamic Block-Sparse Attention
浏览论文内容
中文总结 AI 辅助
针对视频扩散模型中动态块稀疏注意力因逻辑掩码与GPU执行耦合而效率低的问题,提出Tessera运行时,通过物理映射与任务组织解耦二者,结合配置引导选择,实现高达6.79倍加速。
中文摘要 AI 辅助
注意力计算使得视频扩散变换器(vDiTs)中的推理成本高昂,这些变换器通过迭代去噪生成视频。块稀疏注意力(BSA)通过仅计算由逻辑掩码选定的块来降低这一成本,逻辑掩码指定了要计算的注意力交互。然而,将逻辑块几何形状与执行选择耦合在一起,限制了对不同掩码和图形处理单元(GPU)的适应性,而运行时内核特化可能产生准备开销,其成本超过执行时间节省。我们提出了Tessera,一个专为动态BSA设计的专用运行时,它在保留指定注意力交互的同时,将逻辑掩码与GPU执行解耦。其物理映射层将逻辑注意力块保留、合并或细分为适合不同注意力掩码形状和GPU架构的物理瓦片。其任务组织层在GPU任务内对瓦片进行分组和调度,以重用数据、暴露并行性,并将数据移动与计算重叠。最后,基于配置文件的机制选择通过由离线分析构建的查找表实现低开销的执行计划选择。我们使用支持四代NVIDIA GPU的专用CUDA内核实现了Tessera。在2,315个真实注意力掩码和工业视频扩散模型上评估,Tessera在所评估的视频扩散模型中实现了高达6.79倍的BSA请求加速比,相对于基线系统。
英文摘要
Attention computation makes inference expensive in video diffusion transformers (vDiTs), which generate videos through iterative denoising. Block-sparse attention (BSA) reduces this cost by computing only blocks selected by a logical mask, which specifies attention interactions to compute. However, coupling logical block geometry to execution choices limits adaptation to varying masks and graphics processing units (GPUs), while runtime kernel specialization can incur preparation overhead that outweighs execution time savings. We present Tessera, a specialized runtime for dynamic BSA that decouples logical masks from GPU execution while preserving specified attention interactions. Its physical mapping layer retains, combines, or subdivides logical attention blocks into physical tiles suited to different attention mask shapes and GPU architectures. Its task organization layer groups and schedules tiles within GPU tasks to reuse data, expose parallelism, and overlap data movement with computation. Finally, profile-guided regime selection enables low- overhead execution plan selection through a lookup table constructed from offline profiling. We implement Tessera with specialized CUDA kernels supporting four NVIDIA GPU generations. Evaluated on 2,315 real attention masks and industrial video diffusion models, Tessera achieves up to 6.79x BSA request speedup over baseline systems in the evaluated video diffusion models.
发表机构
- National University of Singapore(新加坡国立大学)
机构由 AI 辅助整理,请以论文原文为准。