发表机构
Tel Aviv University(特拉维夫大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出改进的查询无关softmax注意力coreset构造,通过球面提升和Chevet不等式将大小降至O(e^ρ(√d_v+√d_k√log(1+ρ))/ε),并给出完整证明及通信下界转移。
AI 中文摘要
针对softmax注意力头的一个查询无关的coreset是键值对的一个子集S,使得仅由S计算的注意力在ℓ2范数下与完整输出的差距在ε以内,并且对于球内的每一个查询同时成立。Liberty、Andoni和Kleiner证明了大小为O(√d e^(ρ+1/2 log ρ+o(log log ρ))/ε)的无权重coreset存在,其中ρ是查询半径乘以中心化键半径,而下界为Ω(√d e^ρ/ε),他们猜想缩小这一差距需要新技术。我们证明并非如此。将两个球进行球面提升到一个指数核实例中,使得Bozzai和Rothvoss的链式界可以直接应用,并且Chevet不等式将键维度和值维度分开:大小为O(e^ρ(√d_v+√d_k√log(1+ρ))/ε)的无权重coreset存在,并且可以在随机多项式时间内计算,这是第一个在存在性大小上(直到√log(1+ρ)因子)具有整个球保证的结果。一个与维度无关的采样上限O(e^(2ρ)/ε²)完善了这一框架。在固定维度下,对数因子消失:将键球补全为球面使得核成为无权重高斯核,因此Tai的与直径无关的界给出O_{d_k,d_v}(e^ρ/ε),排除了该情形下匹配的对数下界,并回答了Bozzai和Rothvoss关于指数核和Hellinger核的问题在高斯限制情形下的情况。我们以中心化约定重新陈述Liberty–Andoni–Kleiner界并给出完整证明,并表明Chen等人的单向通信界可以转移到查询无关的coreset,其中对于ε≪e^{-ρ},这些界是已知的最强下界。维度因子是所有查询进行一次签名的代价:对于单个查询,差异为O(e^ρ),与维度无关。
英文摘要
A query-oblivious coreset for a softmax-attention head is a subset of the key-value pairs whose attention output is within $\varepsilon$ of the full one for every query in a ball. Liberty, Andoni and Kleiner proved that unweighted coresets of size $O(\sqrt d e^{ρ+\frac12\logρ+o(\log\logρ)}/\varepsilon)$ exist, $ρ$ the query radius times the centred key radius, against a lower bound $Ω(\sqrt d e^ρ/\varepsilon)$, and conjectured that closing the gap needs new techniques. It does not: a spherical lift of both balls into one exponential-kernel instance lets the Bozzai-Rothvoss chaining bound apply, and Chevet's inequality splits key from value dimension, giving coresets of size $O(e^ρ(\sqrt{d_v}+\sqrt{d_k\log(1+ρ)})/\varepsilon)$ in randomised polynomial time, the first constructive whole-ball guarantee within $\sqrt{\log(1+ρ)}$ of the lower bound. A sampling cap $O(e^{2ρ}/\varepsilon^{2})$ completes the envelope; in fixed dimension Tai's diameter-free bound removes the logarithm, settling the Gaussian-restriction case of a Bozzai-Rothvoss question for the kernels. We give theLiberty-Andoni-Kleiner lower boud transfer the one-waycommunication bounds of Chen et r is the price of one signing forall queries. A census of every head of Qwen2.5-7B-Instruct and Llama-3-8B-Instruct finds $ρ$ at least 23.877, so everyactor $e^ρ/\varepsilon$prescribes a coreset larger than the cache: the algorithmic contribution is asymptotic on these models.