arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.35629cs.LGcs.CL

SANTA++:通过代表性键进行采样注意力

SANTA++: Sampling Attention through Representative Keys

Kyle Lee, Christian Z. Pratt, Ruoyu Fang, Heekyung Lee, Avinash Lohitsa, Ryan Modafe, Kerem Y. Camsari

首次发表
浏览论文内容

中文总结 AI 辅助

SANTA++是一种无需训练的随机注意力方法,通过代表性键采样团队并利用重要性采样校正,在减少KV读取的同时保持高注意力准确性,并在长上下文场景下实现显著加速。

中文摘要 AI 辅助

注意力机制通常集中在上下文中的一小部分令牌上,但具体是哪一部分会因查询而异。为了利用这种变化的结构,我们引入了SANTA++,一种无需训练的随机注意力方法,它使用代表性键进行内存高效的筛选,而无需扫描整个键值(KV)缓存。缓存中的键被组织成团队,查询对每个团队的一个代表进行评分,以决定采样哪些团队。我们在采样的团队内计算精确的注意力分数,并通过包含概率的倒数对每个团队的贡献进行重新加权。这种重要性采样校正估计了完整缓存上的注意力,其采样预算让我们可以在内存读取和准确性之间进行权衡。值得注意的是,在32或64个采样团队的情况下,SANTA++使用了密集注意力KV读取的16%至22%,并在LongBench v2和HELMET的检索增强生成子集上保留了密集注意力基线分数的94%至99%,在RULER上保留了85%至91%,使用Qwen2.5-7B-Instruct在32K上下文下。在31个采样团队的情况下,我们的GPU实现在32K上下文下比密集FlashAttention基线实现了1.69倍的注意力加速。通过减少读取的缓存条目数量,SANTA++原则上可以补充具有压缩KV表示的架构,例如多头潜在注意力。我们的内核可在以下网址获取:this https URL。

英文摘要

Attention often concentrates on a small subset of tokens in the context, but which subset matters changes from one query to the next. To exploit this changing structure, we introduce SANTA++, a training-free stochastic attention method that uses representative keys for memory-efficient selection without scanning the entire key-value (KV) cache. Cached keys are organized into teams, and the query scores one representative from each team to decide which teams to sample. We compute exact attention scores within the sampled teams and reweight each team's contribution by the inverse of its inclusion probability. This importance sampling correction estimates attention over the full cache, with a sampling budget that lets us trade memory reads for accuracy. Remarkably, with 32 or 64 sampled teams, SANTA++ uses 16% to 22% of dense attention's KV reads and retains 94% to 99% of the dense-attention baseline's scores on LongBench v2 and HELMET's retrieval-augmented generation subset, and 85% to 91% on RULER, with Qwen2.5-7B-Instruct at 32K context. With 31 sampled teams, our GPU implementation delivers a $1.69\times$ attention speedup over the dense FlashAttention baseline at 32K context. By reducing the number of cache entries read, SANTA++ in principle complements architectures with compressed KV representations, such as multi-head latent attention. Our kernels are available at: https://github.com/OPUSLab/santapp-kernel-demo.git.

发表机构

  • University of California, Santa Barbara(加州大学圣塔芭芭拉分校)
  • Flucta

机构由 AI 辅助整理,请以论文原文为准。

↑