arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

通过黑盒向量搜索实现注意力机制

Attention via Black-Box Vector Search

Stepan Zharkov, Krish Singal, Ashwin Padaki, Alexandr Andoni

arXiv 2610.10135首次发表:更新:

发表机构

Columbia University; University of Pennsylvania(哥伦比亚大学; 宾夕法尼亚大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文通过优先级采样框架研究黑盒向量搜索下的稀疏注意力估计,证明单索引需检索Θ(√n/ε)个键,多索引可降至O(log n+1/ε²),并设计增强方法绕过下界,在LLM推理中优于现有方法。

AI 中文摘要

稀疏注意力机制使用一小部分键子集来估计对 $n$ 个标记的注意力。许多现有方法使用最大内积搜索(MIPS)来检索最重的键,这引出了以下问题:给定对 MIPS 预言机的黑盒访问,必须检索多少个键才能输出一个 $\varepsilon$-精确的注意力估计?我们通过优先级采样的框架统一了先前的方法来回答这个问题。使用单个 MIPS 索引,我们证明检索 $\Theta(\sqrt{n}/\varepsilon)$ 个键既是充分的也是必要的。使用 $\Theta(\log n)$ 个索引,我们给出了一种仅检索 $O(\log n+1/\varepsilon^2)$ 个键的算法,并证明这是近乎最优的。更一般地,我们设计了在 MIPS 索引数量和检索键数量之间建立平滑权衡的算法。然后我们表明,如果允许对键和查询进行增强,我们可以绕过上述下界:存在一个使用单个 MIPS 索引和 $O(1/\varepsilon^2)$ 个检索键的简单优先级采样估计器。当集成到 LLM 推理中时,我们的算法优于先前工作中使用的 top-$k$ 和采样方法,并产生能良好扩展到长上下文的注意力近似。

英文摘要

Sparse attention mechanisms estimate attention over $n$ tokens using a small subset of keys. Many existing approaches use maximum inner product search (MIPS) to retrieve the heaviest keys, which motivates the following question: given black-box access to a MIPS oracle, how many keys must be retrieved to output an $\varepsilon$-accurate attention estimate? We answer this question by unifying prior approaches through the framework of priority sampling. With a single MIPS index, we show that $Θ(\sqrt{n}/\varepsilon)$ retrieved keys are both sufficient and necessary. With $Θ(\log n)$ indices, we give an algorithm that retrieves only $O(\log n+1/\varepsilon^2)$ keys and prove that this is near-optimal. More generally, we design algorithms that establish a smooth tradeoff between the number of MIPS indices and number of retrieved keys. We then show that if we allow augmentation of keys and queries, we can bypass the above lower bounds: there exists a simple priority-sampling estimator using a single MIPS index and $O(1/\varepsilon^2)$ retrieved keys. When integrated into LLM inference, our algorithms outperform top-$k$ and sampling approaches used in prior work and yield attention approximation that scales favorably to long contexts.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑