arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Page-EntroKV:分组查询注意力下硬件对齐的熵加权KV缓存驱逐

Page-EntroKV: Hardware-Aligned, Entropy-Weighted KV-Cache Eviction under Grouped-Query Attention

Inbasekaran S

arXiv 2610.03135首次发表:更新:

发表机构

SRM Institute of Science and Technology(SRM科学与技术学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对GQA下KV缓存驱逐的硬件对齐问题,提出Page-EntroKV框架,利用熵加权池化与页级驱逐,消除并集开销,提升针检索率并保持预算精确。

AI 中文摘要

服务长上下文自回归语言模型受到键值(KV)缓存的限制。大多数动态驱逐方法按查询头对令牌重要性进行评分,并独立选择令牌。这与分组查询注意力(GQA)的适配性不佳,在GQA中,多个查询头共享一个物理KV缓冲区:发散的各头选择迫使服务引擎保留其选择的并集——这会使缓存膨胀高达组比率r——而算术平均池化会稀释承担事实回忆的专门检索头。我们引入Page-EntroKV,一个在GQA服务实际分配的粒度上运行的KV缓存驱逐的正式框架。每个物理组内的头通过由孤立于汇聚(sink)的碰撞(Renyi-2)熵导出的权重进行池化——每个头一个内积,在预填充时计算一次,无需校准——因此汇聚头不能伪装成检索头。池化分数被投影到PagedAttention页帧上,驱逐在硬件元组(层、组、页)上执行。我们形式化了并集开销比率(UOR)和组内不一致性,证明了一个将二者联系起来的精确恒等式(适用于两头组)以及每个组比率下的两侧界限,证明了严格的预算保持和算术平均池化可证明违反的有限上下文针检索界限,并给出了精确的逐层页记账。在一个试点架构(Qwen2.5-1.5B-Instruct,r=6)上,对2240个组测量进行头独立重放,在2%预算下产生高达4.75倍的并集开销,而Page-EntroKV将UOR精确保持在1.000;汇聚隔离消除了13倍的汇聚伪装;在20%预算下,针检索率为100%,而平均池化为0%;保留基数对每个页大小都是精确的;QA和代码任务在20%保留率下仍可解决。

英文摘要

Serving long-context autoregressive language models is constrained by the key-value (KV) cache. Most dynamic eviction methods score token importance per query head and choose tokens independently. This fits poorly with grouped-query attention (GQA), where several query heads share one physical KV buffer: divergent per-head selections force the serving engine to retain the union of their choices - inflating the cache by up to the group ratio r - while arithmetic mean pooling dilutes the specialized retrieval heads that carry factual recall. We introduce Page-EntroKV, a formal framework for KV-cache eviction operating at the granularity GQA serving actually allocates. Heads within each physical group are pooled by weights derived from sink-isolated collision (Renyi-2) entropy - one inner product per head, computed once at prefill with no calibration - so sink heads cannot masquerade as retrieval heads. Pooled scores are projected onto PagedAttention page frames, and eviction executes at the hardware tuple (layer, group, page). We formalize the union overhead ratio (UOR) and intra-group disagreement, prove an exact identity linking them for two-head groups alongside two-sided bounds at every group ratio, prove strict budget preservation and a finite-context needle-retention bound that arithmetic mean pooling provably violates, and give exact per-layer page accounting. On a pilot architecture (Qwen2.5-1.5B-Instruct, r=6), head-independent replay over 2,240 group measurements yields union overhead up to 4.75x at a 2% budget, while Page-EntroKV holds UOR exactly 1.000; sink isolation removes a 13x sink masquerade; needle recall is 100% versus 0% for mean pooling at a 20% budget; retained cardinality is exact for every page size; and QA and code tasks remain solvable at 20% retention.

Comments24 pages, 8 figures, 10 tables. Formal framework with pilot-scale empirical validation on Qwen2.5-1.5B-Instruct. Includes step-by-step derivations (Appendix C) and PyTorch reference implementation (Appendix D). Code and data available at: https://github.com/bruce12-glitch/PageEntro-KV

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑