发表机构
Technion; Crusoe AI; Corma; Stealth Startup(以色列理工学院; Crusoe AI; Corma; Stealth Startup)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出SAGA非对称稀疏注意力,通过解耦键头与值头数量并配合Atop-N近似稀疏注意力,在长上下文中实现超2倍解码加速,且从头训练模型质量接近GQA基线,并提供高效微调方法便于采用。
AI 中文摘要
大型语言模型(LLM)中的自回归生成受到注意力机制的内存和计算需求的限制。稀疏注意力方法通过仅选择注意力矩阵中高概率的条目来缓解这一成本。我们观察到,在许多此类方法中,这使得概率-值乘法变得可忽略,从而将瓶颈转移到查询-键步骤。因此,可以减少键头以加速推理,而保留更多的值头可以在有限的额外解码成本下保持容量。我们引入了稀疏非对称分组查询注意力(SAGA),它解耦了键头和值头的数量以利用这一原理,并将其与近似Top-N(Atop-N)注意力配对,这是一种简单的稀疏注意力方法,旨在研究稀疏性与头数不对称之间的相互作用。我们在理论上形式化了这种不对称性的好处,并通过在多达15亿参数的模型上的延迟测量和质量评估进行了实证验证。SAGA和Atop-N共同在长上下文中实现了比我们的全注意力GQA基线超过2倍的端到端解码加速。从头训练的SAGA模型在评估基准上几乎与可比的GQA变体质量相当。为了便于采用,我们引入了一种高效的微调方法,将预训练模型转换为SAGA架构,使从业者无需昂贵的重新训练即可受益于我们的方法。
英文摘要
Autoregressive generation in Large Language Models (LLMs) is constrained by the memory and computational demands of attention mechanisms. Sparse attention methods mitigate this cost by selecting only high-probability entries of the attention matrix. We observe that in many such methods, this renders the probability-value multiplication negligible, shifting the bottleneck to the query-key step. Key heads can therefore be reduced to accelerate inference, while retaining more value heads preserves capacity with limited additional decoding cost. We introduce Sparse Asymmetric Group-Query Attention (SAGA), which decouples key and value head counts to exploit this principle, and pair it with approximate top-N (Atop-N) attention, a simple sparse attention method designed to study the interaction between sparsity and head-count asymmetry. We formalize the benefits of this asymmetry theoretically and validate them empirically through latency measurements and quality evaluations on models up to 1.5B parameters. Together, SAGA and Atop-N achieve end-to-end decoding speedups exceeding $2\times$ over our full-attention GQA baseline at long contexts. Models trained from scratch with SAGA nearly match the quality of comparable GQA variants on the evaluated benchmarks. To facilitate adoption, we introduce an efficient fine-tuning method that converts pretrained models to the SAGA architecture, enabling practitioners to benefit from our approach without costly retraining.
CommentsAccepted to NeurIPS 2026