AVQ注意力:自适应向量量化注意力
AVQ-Attention: Adaptive Vector-Quantized Attention
浏览论文内容
中文总结 AI 辅助
研究针对Transformer模型注意力机制计算瓶颈,提出自适应向量量化(AVQ)注意力,基于注意力重要性分配码本容量,前向传播识别重要代码并用子码字细化,开发相关实现,保持\(\mathcal{O}(MN)\)复杂度,提升精度-效率权衡。
中文摘要 AI 辅助
注意力机制在N个令牌上的\(\mathcal{O}(N^2)\)复杂度仍是Transformer模型的计算瓶颈。向量量化(VQ)注意力通过用M个码字表示键将其降至\(\mathcal{O}(MN)\),但无论注意力集中何处都采用均匀码本容量。我们提出自适应向量量化(AVQ)注意力,基于注意力重要性自适应分配码本容量。从少量码字开始,前向传播中识别最重要代码并用预学习子码字细化,在关键处实现细粒度量化,其他处保持粗量化。开发了基于自定义Triton内核的实现,在Flash Attention的平铺计算范式内以最小开销完成完整自适应细化过程。该方法保持\(\mathcal{O}(MN)\)复杂度,与固定码本VQ注意力相比,实现了更好的精度-效率权衡。
英文摘要
The $\mathcal{O}(N^2)$ complexity of attention over $N$ tokens remains a computational bottleneck in transformer models. Vector-Quantized (VQ) attention reduces this to $\mathcal{O}(MN)$ by representing keys with $M$ codewords, but applies uniform codebook capacity regardless of where attention mass concentrates: high-attention regions of key space may be coarsely approximated while low-attention regions waste representational capacity. We propose Adaptive Vector-Quantized (AVQ) Attention, which adaptively allocates codebook capacity based on attention importance. Starting from a small set of codewords, our method identifies the most important codes during the forward pass and refines them with pre-learned child codewords, achieving fine-grained quantization where it matters most while maintaining coarse quantization elsewhere. We develop an implementation using custom Triton kernels that enables the full adaptive refinement process, including importance scoring, child codeword insertion, and parent contribution replacement, to be carried out within the tiled computation paradigm of Flash Attention with minimal overhead. Our approach maintains $\mathcal{O}(MN)$ complexity while achieving improved accuracy-efficiency trade-offs compared to fixed-codebook VQ-attention.
发表机构
- QUVA Lab, University of Amsterdam(QUVA实验室,阿姆斯特丹大学)
- AMLab, Informatics Institute, University of Amsterdam(AMLab,阿姆斯特丹大学信息学院)
- AI4Science Lab, University of Amsterdam(AI4Science实验室,阿姆斯特丹大学)
- Korteweg-de Vries Institute for Mathematics, University of Amsterdam(科特韦格 - 德弗里斯数学研究所,阿姆斯特丹大学)
- Qualcomm AI Research(高通人工智能研究公司)
- FunAI Lab, University of Technology Nuremberg(FunAI实验室,纽伦堡工业大学)
机构由 AI 辅助整理,请以论文原文为准。