arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.20214cs.LGcs.AI

ELSAA:用于训练Transformer的高效低秩和稀疏注意力近似

ELSAA: Efficient Low-Rank and Sparse Attention Approximation for Training Transformers

Mahdi Heidari, Mohammad Mahdi Rahimi, Jaekyun Moon

首次发表
浏览论文内容

中文总结 AI 辅助

研究旨在解决将Transformer扩展到更长输入长度的问题,提出ELSAA方法,通过近似注意力分数算子,结合稀疏与低秩分支及分母感知融合项,构建低秩稀疏注意力输出,实现更长上下文训练并保留交互与混合。

中文摘要 AI 辅助

二次的N×N注意力分数矩阵仍然是将Transformer扩展到更长输入长度的主要障碍。现有高效注意力方法通常通过施加稀疏性(使每个查询仅关注一小部分键)或使用低秩/内核草图(将全局交互压缩到低维表示)来减少此瓶颈。我们提出了ELSAA,一种高效的低秩和稀疏注意力近似。重要的是,ELSAA不会将Transformer的学习投影或输出矩阵分解为稀疏和低秩因子。相反,在密集投影产生Q、K、V后,ELSAA近似诱导注意力分数算子本身:一个稀疏分支捕获选定的高相似性交互,一个低秩分支总结扩散的全局交互。由于两个分支可以在具有非常不同分母质量的支持上进行归一化,ELSAA引入了一个分母感知融合项,根据其相对于低秩分支的估计注意力质量来缩放稀疏分支。这给出了一个实用框架,用于构建低秩和稀疏注意力输出而无需实现完整的二次分数矩阵,旨在实现更长上下文训练,同时保留清晰的令牌级交互和广泛的上下文混合。

英文摘要

The quadratic $N\times N$ attention score matrix remains a central obstacle to extending Transformers to longer input lengths. Existing efficient attention methods usually reduce this bottleneck by either imposing sparsity, so that each query attends to only a small subset of keys, or by using low-rank/kernel sketches, so that global interactions are compressed into a lower-dimensional representation. We propose \emph{ELSAA}, an efficient low-rank and sparse approximation of attention. Importantly, ELSAA does \emph{not} decompose the learned projection or output matrices of the Transformer into sparse and low-rank factors. Instead, after dense projections produce $Q,K,V$, ELSAA approximates the induced attention score operator itself: a sparse branch captures selected high-similarity interactions, while a low-rank branch summarizes diffuse global interactions. Since the two branches can be normalized over supports with very different denominator mass, ELSAA introduces a denominator-aware fusion term that scales the sparse branch according to its estimated attention mass relative to the low-rank branch. This gives a practical framework for constructing low-rank and sparse attention outputs without materializing the full quadratic score matrix, aiming to enable longer-context training while preserving both sharp token-level interactions and broad contextual mixing.

↑