发表机构
Qualcomm AI Research(高通人工智能研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对现代LLM中弱注意力汇聚导致现有KV缓存驱逐失效的问题,提出基于值向量几何分散的ValueDiff驱逐方法,在多个基准上显著提升性能。
AI 中文摘要
现代LLM采用QK归一化、门控注意力、学习到的注意力汇聚或logit softcapping,表现出较弱的持久注意力汇聚,而现有的KV缓存驱逐方法主要依赖这些汇聚。我们观察到,在这些模型中,较弱的汇聚伴随着相对于键向量分散更大的值向量分散。受这种值侧分散的启发,我们提出了ValueDiff,一种值几何驱逐方法,根据值向量与缓存均值的L2偏差对令牌进行排序。在关于未来注意力的最大熵假设下,相同的分数作为最小干扰驱逐出现。我们在固定缓存预算下进行评估,在预填充期间每个块边界以及生成期间每个解码步骤进行驱逐。在RULER上,在严格的2k令牌预算下,ValueDiff在七个抑制汇聚模型中保留了密集的88%至99%(在7个中最佳为6个)。在LongBench的4k预算下,ValueDiff在抑制汇聚模型上平均保留92%,而最强先前基线为83%。在MATH-500上,在25%缓存预算下,ValueDiff在每个测试的抑制汇聚模型上都是最强的非密集方法,在门控注意力模型上优于先前方法高达约20个百分点。在所有三个基准测试中,值几何成为抑制汇聚模型更可靠的查询不变驱逐信号。
英文摘要
Modern LLMs with QK-normalization, gated attention, learned attention sinks, or logit softcapping exhibit weaker persistent attention sinks, on which existing KV cache eviction methods primarily rely. We observe that across these models, weaker sinks co-occur with greater value-vector dispersion relative to key-vector dispersion. Motivated by this value-side dispersion, we present ValueDiff, a value-geometric eviction that ranks tokens by the L2 deviation of their value vectors from the cache mean. The same score arises as the minimal-disturbance eviction under a max-entropy assumption about future attention. We evaluate under fixed cache budgets, with eviction at every block boundary during prefill and at every decoding step during generation. On RULER at a tight 2k token budget, ValueDiff retains 88-99% of dense across seven sink-suppressed models (best on 6 out of 7). On LongBench at the 4k budget, ValueDiff averages 92% retention across sink-suppressed models versus 83% for the strongest prior baseline. On MATH-500, ValueDiff is the strongest non-dense method on every sink-suppressed model tested at the 25% cache budget, outperforming prior methods by up to ~20 points on gated-attention models. Across all three benchmarks, value geometry emerges as the more reliable query-invariant eviction signal for sink-suppressed models.
Comments10 pages, 3 Figues