arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

后端无关的稀疏注意力用于快速高分辨率视觉生成

Backend-Agnostic Sparse Attention for Fast High-Resolution Visual Generation

Liao Ma, Jiayi Song, Yunfeng Wu, Songhua Liu, Peilin Zhao

arXiv 2610.08772首次发表:更新:

发表机构

School of Artificial Intelligence, Shanghai Jiao Tong University; School of Data Science, Fudan University; School of Computing and Data Science, The University of Hong Kong(上海交通大学人工智能学院; 复旦大学数据科学学院; 香港大学计算及数据科学学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对扩散变换器高分辨率生成中全注意力计算昂贵的问题,提出后端无关的稀疏注意力 BASA,通过跨块窗口移位实现全局交互,无需定制内核,在 FLUX 和 Wan 上实现接近理论值的加速并保持生成质量。

AI 中文摘要

扩散变换器(DiTs)在图像和视频生成中取得了强劲的性能,但全注意力的二次复杂度使得高分辨率生成在计算上代价高昂。窗口注意力提供了一种高效的替代方案,然而现有方法面临一个实际的权衡:分区窗口注意力通常能达到与其理论复杂度一致的计算效率。然而,孤立的窗口阻断了跨窗口的交互,往往在生成结果中引入可见的网格状伪影。细粒度的滑动窗口注意力有效地恢复了相邻窗口间的交互并提升了视觉质量。然而,其不规则的 computation 模式造成了理论加速与实际加速之间的显著差距,并且需要针对每个硬件后端定制专门的内核。为了解决这些挑战,我们提出了 BASA,一种后端无关的稀疏注意力,它兼具两者的优点:视觉质量与实际加速。具体来说,BASA 用移位的局部窗口注意力替代视觉自注意力。通过在 DiT 块之间引入结构化的窗口移位方案,我们允许在一层中被窗口边界分割的 token 在后续层中进行通信,从而实现全局信息交换并消除窗口引起的视觉伪影。值得注意的是,我们的设计不引入额外的不规则算子或定制内核,使其能够轻松部署在现有的注意力后端上,并弥合了理论稀疏性与实际加速之间的差距。实验表明,BASA 在 FLUX 上实现了超过理论估计 90% 的实测加速,并在 Wan 上提供了 4.52 倍的注意力加速,同时保持了有竞争力的生成质量。

英文摘要

Diffusion Transformers (DiTs) have achieved strong performance in image and video generation, but the quadratic complexity of full attention makes high-resolution generation computationally expensive. Window attention offers an efficient alternative, yet existing methods face a practical trade-off: partitioned window attention typically achieves computational efficiency consistent with its theoretical complexity. However, isolated windows block cross-window interaction, often introducing visible grid-like artifacts in the generated results. Fine-grained sliding-window attention effectively restores interactions across neighboring windows and improves visual quality. However, its irregular computation patterns create a substantial gap between theoretical and practical speedups and require specialized kernels tailored to each hardware backend. To tackle these challenges, we propose BASA, a backend-agnostic sparse attention, which brings the best of both worlds: visual quality and practical acceleration. Specifically, BASA replaces visual self-attention with shifted local-window attention. By introducing a structured window-shifting scheme across DiT blocks, we allow tokens divided by window boundaries in one layer to communicate in the following layers, thereby achieving global information exchange and eliminating window-induced visual artifacts. Notably, our design introduces no additional irregular operators or customized kernels, making it readily deployable on existing attention backends and closing the gap between theoretical sparsity and practical acceleration. Experiments demonstrate that BASA achieves measured speedups exceeding 90\% of the theoretical estimates on FLUX and delivers a 4.52$\times$ attention speedup on Wan while maintaining competitive generation quality. Codes are publicly available at: https://github.com/lama0110/BASA.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑