arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

超越均值注意力:面向KV缓存驱逐的多样性感知分层评分

Beyond Mean Attention: Diversity-Aware, Layer-Wise Scoring for KV Cache Eviction

Tianfang Xie, Wei Zhu

arXiv 2609.30738首次发表:更新:

发表机构

Georgia Institute of Technology; Zhangjiang Lab(佐治亚理工学院; 张江实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对KV缓存驱逐,提出统一评分(均值+离散度+冗余惩罚),验证分层系数在LongBench上多数数据集有效,段落检索中中间层符号翻转带来显著增益。

AI 中文摘要

KV缓存驱逐方法(如SnapKV和PyramidKV)仅通过小观察窗口上的均值注意力对令牌进行排序。我们研究了一个统一评分,$\mu_i+\lambda_1\sigma_i+\lambda_2\mathrm{corr}(i,S)$,其中增加了跨窗口查询的注意力离散度以及相对于已选令牌的冗余度。对于$\lambda_2<0$,该评分会像最大边际相关性(MMR)一样惩罚与已选令牌的相似性,而无需额外的前向传播。为了测试这种相关性与多样性之间的平衡是否应随深度变化,我们将固定的全局系数与三段式和二次型配置文件进行了比较。仅在开发集上通过$\sinh$重新参数化搜索这些深度配置文件。在所有16个英文LongBench数据集上,使用Mistral-7B且每层预算为64个条目时,单一的全局多样化常数改善了16个数据集中的13个(宏观+1.1);该增益在预算为32时保持,在128时缩小。按数据集搜索在大多数数据集上未发现可检测的层结构;在段落检索上发现了一个大的层结构:一个中间层的符号翻转,奖励相似性,在预算为64时比基线高+9.6,在预算为128时无需重新调整即比全局常数高+13.2。消融实验将该增益归因于冗余项;在保留的测试集上重放每个接受的搜索状态,将真正的结构与调参噪声区分开来。

英文摘要

KV cache eviction methods such as SnapKV and PyramidKV rank tokens solely by mean attention over a small observation window. We study a unified score, $μ_i+λ_1σ_i+λ_2\mathrm{corr}(i,S)$, adding attention dispersion across window queries and redundancy relative to selected tokens. For $λ_2<0$, the score penalizes similarity to selected tokens as in maximal marginal relevance (MMR), without extra forward passes. To test whether this relevance-diversity balance should vary with depth, we compare fixed global coefficients with three-segment and quadratic profiles. Only these depth profiles are searched on a development split under a $\sinh$ reparameterization. On all 16 English LongBench datasets with Mistral-7B at a budget of 64 entries per layer, a single global diversification constant improves 13 of 16 datasets (macro +1.1); the gain holds at budget 32 and narrows at 128. Per-dataset search finds no detectable layer structure on most datasets; on passage retrieval it finds a large one: a mid-layer sign flip that rewards similarity and is worth +9.6 over the baseline at budget 64 and, without re-tuning, +13.2 over the global constant at budget 128. Ablations attribute the gain to the redundancy term; replaying every accepted search state on the held-out test set separates genuine structure from tuning noise.

Comments5 pages, 1 figure, 3 tables. Submitted to IEEE ICASSP 2027

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑