arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.10871cs.CL

稀疏注意力是矩阵近似,而非从一组值中进行选择

Sparse Attention Is Matrix Approximation, Not Choosing from a Bag of Values

  • Stony Brook University(石溪大学)
  • Duke University(杜克大学)

机构由 AI 辅助整理,请以论文原文为准。

Fang Wan, Xufeng Liu, Fan Li, Yi Liu

中文总结 AI 辅助

该研究指出现有稀疏注意力方法的核心概念错误,提出矩阵近似稀疏注意力(MASA),将其作为插件式修正加入现有框架,经多组实验验证可提升准确率,支持稀疏注意力的矩阵近似观点。

中文摘要 AI 辅助

大型语言模型(LLMs)在诸多领域表现出色,但注意力机制随提示长度呈二次方增长的成本限制了其效率。稀疏注意力通过仅保留小部分查询-键交互来近似完整注意力矩阵,以此降低成本。然而,现有方法陷入了数学上错误的观点:它们仅保留注意力矩阵的大标量条目或高质量区域,将注意力矩阵视为一组值,忽略了它是结构化矩阵,其条目会通过与值向量相乘共同决定注意力输出。我们认为这是核心概念问题:稀疏注意力应被表述为矩阵近似,而非盲目从一组条目中选择最大值。基于此观点,我们提出矩阵近似稀疏注意力(Matrix Approximation Sparse Attention,MASA)。MASA用闭式评分取代原始注意力质量排序,该评分衡量每个稀疏单元减少矩阵乘积近似误差的程度。作为基于理论的插件式修正,MASA可添加到现有稀疏注意力框架中,无需改变其稀疏核或预算。在多种稀疏注意力方法、基准及LLM主干上开展的大量实验显示,其准确率持续提升,既验证了MASA的有效性,也支持稀疏注意力的矩阵近似观点。

英文摘要

Large Language Models (LLMs) achieve strong performance across many domains, but their efficiency is limited by the quadratic cost of attention with respect to prompt length. Sparse attention reduces this cost by retaining only a small fraction of query-key interactions to approximate the full attention matrix. However, existing methods are trapped in a mathematically wrong view: they simply keep large scalar entries or high-mass regions of the attention matrix. This treats the attention matrix as a bag of values, ignoring that it is used as a structured matrix whose entries jointly determine the attention output through multiplication with value vectors. We argue that this is the core conceptual issue: sparse attention should be formulated as matrix approximation, not as blindly choosing the largest values from a bag of entries. Based on this view, we propose Matrix Approximation Sparse Attention (MASA). MASA replaces raw attention-mass ranking with a closed-form score that measures how much each sparse unit reduces matrix-product approximation error. As a theory-grounded plug-in correction, MASA can be added to existing sparse attention frameworks without changing their sparse kernels or budgets. Extensive experiments across multiple sparse attention methods, benchmarks, and LLM backbones show consistent accuracy gains, supporting both MASA and the matrix-approximation view of sparse attention.

↑