arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

理解并利用自回归图像生成中的对角注意力稀疏性

Understanding and Exploiting Diagonal Attention Sparsity in Autoregressive Image Generation

Daeun Kim, Junwha Hong, Changhun Oh, Yoonsung Kim, Yoonhyeong Lee, Jongse Park

arXiv 2609.19702首次发表:更新:

发表机构

KAIST; Agency for Defense Development; Seoul National University(韩国科学技术院; 国防科学研究所; 首尔大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究首次系统表征自回归图像生成中的注意力稀疏性,发现对角稀疏模式,并提出对角感知稀疏注意力机制,实现高达3.1倍吞吐量提升和1.19倍延迟改善,质量损失小于2%。

AI 中文摘要

自回归图像生成因其与基于Transformer的LLM服务基础设施的兼容性,已成为多模态AI系统的一种范式。然而,每个请求生成数千个视觉标记,使得解码过程中注意力计算时的KV缓存访问日益成为瓶颈。稀疏注意力对此工作负载特别有吸引力,因为许多视觉生成应用容忍适度质量下降以换取性能和效率的提升。尽管稀疏注意力已在基于文本的LLM推理中得到广泛探索,但其稀疏性假设是否能有效推广到自回归图像生成仍不清楚。我们首次对自回归图像生成中跨多样工作负载和代表性开源模型的注意力稀疏性进行了系统性表征。我们的分析揭示了几个显著特性,包括明显的前缀-解码不对称性、对提示和局部标记的强注意力集中,以及由视觉标记的空间局部性产生的独特对角注意力稀疏模式。受这些观察启发,我们提出了一种对角感知的稀疏注意力机制,该机制在最近窗口内沿对角注意力方向选择性地跳过KV条目。在使用FlexGen、FlashAttention-2和自定义内核的基于GPU的服务系统上实现后,与密集推理相比,我们的方法实现了高达3.1倍的吞吐量提升和1.19倍的延迟改善,且质量下降不到2%。

英文摘要

Autoregressive image generation has emerged as a paradigm for multimodal AI systems due to its compatibility with transformer-based LLM serving infrastructures. However, generating thousands of visual tokens per request makes decoding increasingly bottlenecked by KV cache accesses during attention computation. Sparse attention is particularly attractive for this workload because many visual generation applications tolerate moderate quality degradation in exchange for improved performance and efficiency. While sparse attention has been extensively explored for text-based LLM inference, it remains unclear whether its sparsity assumptions generalize effectively to autoregressive image generation. We present the first systematic characterization of attention sparsity in autoregressive image generation across diverse workloads and representative open-source models. Our analysis reveals several distinguishing properties, including a pronounced prefill-decode asymmetry, strong attention concentration on prompt and local tokens, and a unique diagonal attention sparsity pattern arising from the spatial locality of visual tokens. Motivated by these observations, we propose a diagonal-aware sparse attention mechanism that selectively skips KV entries along the diagonal attention direction within a recent window. Implemented on top of a GPU-based serving system using FlexGen, FlashAttention-2, and custom kernels, our approach achieves up to 3.1x throughput and 1.19x latency improvements with less than 2% quality degradation compared to dense inference.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑