arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

COBS:累积阶数块稀疏注意力

COBS: Cumulant Order Block Sparse Attention

Alexander Tian, Aditya Ghai, Sanjit Neelam, Zaal Vasania, Akshay Mishra

arXiv 2607.09052首次发表:更新:

发表机构

MatX(MatX)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究大语言模型中缓解KV缓存读取瓶颈的块稀疏注意力方法,以DeepSeek的NSA为代表分析其块选择问题。通过累积量展开指出现有方法局限,提出COBS方法,在32k RULER基准上提升了NSA基线分数,缩小与密集注意力差距,且减少了KV缓存读取流量。

AI 中文摘要

块稀疏注意力是一种硬件友好的方式,可缓解大语言模型中的键值(KV)缓存读取瓶颈。然而,它在领先的开放权重LLMs中并不普遍,这些模型依赖于密集注意力或细粒度选择,从而促使我们进行分析。我们研究了DeepSeek的原生稀疏注意力(NSA)作为一种代表性方法,其三分支设计使我们能够分离出块选择,这是最具挑战性和决定性的阶段。我们将选择形式化,并将其简化为按单个数量对块进行排序,即注意力质量:块的注意力分数之和。我们表明,如果选择检索具有最大注意力质量的块,块稀疏注意力可以匹配密集注意力的质量。然而,计算精确的注意力质量需要读取每个键,因此块选择问题最终归结为从紧凑摘要而不是完整键中近似此质量。通过累积量展开,我们展示了现有方法为何失败:它们的选择策略试图估计注意力质量,但仅限于一阶近似。因此,我们提出了COBS(累积阶数块稀疏注意力),一种基于NSA的注意力方法,结合了一种新颖的选择器,该选择器为每个块存储压缩的二阶统计量。在32k RULER长上下文检索基准上,COBS将NSA基线的平均分数从0.2999提高到0.8195,接近密集注意力的0.9040,缩小了约86%的差距,同时仅使用NSA基线1.21倍的KV缓存读取流量,比密集注意力少1 / 15.15倍的读取流量。在我们的比较中,相同模型保留了短上下文行为,并且比密集注意力具有更低的位置负对数似然(NLL)。

英文摘要

Block sparse attention is a hardware friendly way to alleviate the key-value (KV) cache read bottleneck in large language models (LLMs). However, it is not prevalent among leading open-weight LLMs, which rely instead on dense attention or fine-grained selection, thereby motivating our analysis. We study DeepSeek's Native Sparse Attention (NSA) as a representative method, whose three-branch design lets us isolate block selection, the most challenging and consequential stage. We formalize selection and reduce it to ranking blocks by a single quantity, the attention mass: the sum of a block's attention scores. We show that if selection retrieves the blocks with the largest attention mass, block sparse attention can match the quality of dense attention. However, computing the exact attention mass requires reading every key, so the problem of block selection ultimately reduces to approximating this mass from a compact summary instead of the full keys. Via a cumulant expansion, we show why existing methods falter: their selection strategies attempt to estimate the attention mass, but are confined to a first-order approximation. Therefore, we propose COBS (Cumulant Order Block Sparse Attention), an attention method that builds on NSA, incorporating a novel selector that stores a compressed second-order statistic per block. On the 32k RULER long-context retrieval benchmark, COBS raises the NSA baseline's mean score from 0.2999 to 0.8195, approaching dense attention at 0.9040 and closing about 86% of the gap, while using only 1.21x the KV cache read traffic of the NSA baseline and 15.15x less read traffic than dense. The same model preserves short-context behavior and attains lower position-wise negative log-likelihood (NLL) than dense attention in our comparison.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑