arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.07001cs.LG

每个缓存条目都有其位置:KV缓存压缩的分辨率与覆盖范围的全局分配

Every Cache Entry Earns Its Place: Global Allocation of Resolution and Coverage for KV Cache Compression

Haolin Tian, Yuzhe Liu, Tonghan Wang

首次发表
浏览论文内容

中文总结 AI 辅助

针对LLM长上下文处理的KV缓存瓶颈,提出无需额外训练的GraceKV方法,通过全局资源分配平衡缓存的分辨率与覆盖范围,在32项实验设置中24项排名第一,128倍压缩仍稳健。

中文摘要 AI 辅助

随着大型语言模型(LLMs)处理的上下文长度不断增加,KV缓存的存储与重复访问已成为主要瓶颈。现有的KV缓存压缩方法依赖预定义的固定压缩规则,通常围绕令牌驱逐或合并开发,导致缓存资源无法跨层、跨注意力头和上下文槽自由流动,也无法联合分配以平衡局部分辨率与信息覆盖范围。因此,我们提出GraceKV,一种用于KV缓存压缩中分辨率与覆盖范围分配的全局方法,将压缩过程表述为固定缓存预算下的全局资源分配问题。GraceKV将每个层-KV头-槽组合视为原子单元,并构建原型树:叶节点对应令牌级KV条目,每个内部节点用单个原型压缩其后代覆盖的KV空间;树中一组不重叠的节点构成原子单元的表示,添加新树的根节点可扩展信息覆盖范围,拆分选中节点可提升局部分辨率,所有候选操作在共享缓存预算下全局竞争,最终保留的节点组成压缩后的KV缓存。该过程自适应确定缓存资源在原子单元间的全局分配,以及分辨率与覆盖范围的平衡。GraceKV无需额外训练,整个压缩与推理过程在GPU上执行。在多种长上下文任务和压缩率下的系统实验显示,GraceKV在32个设置中的24个排名第一,在高达128倍压缩时仍保持稳健,这些结果验证了全局预算分配在协调信息覆盖与局部分辨率方面的有效性。

英文摘要

As large language models (LLMs) process increasingly long contexts, KV cache storage and repeated access have become a major bottleneck. Existing KV cache compression methods rely on predefined, fixed compression rules and are typically developed around either token eviction or merging. As a result, cache resources can neither flow freely across layers, heads, and context slots, nor be jointly allocated to balance local resolution and information coverage. Therefore, we propose GraceKV, a global approach for the allocation of resolution and coverage in KV cache compression, and formulate the compression process as a global resource allocation problem under a fixed cache budget. GraceKV treats each layer-KV head-slot combination as an atomic unit and builds a prototype tree. Leaf nodes correspond to token-level KV entries, while each internal node uses a single prototype to compress the KV space covered by its children. A set of non-overlapping nodes in the tree forms the representation of an atomic unit. Adding the root of a new tree expands information coverage, whereas splitting a selected node improves local resolution. All candidate actions compete globally for a shared cache budget. Finally, the nodes retained across all trees form the compressed KV cache. This process adaptively determines the allocation of cache resources among atomic units globally and the balance between resolution and coverage. GraceKV requires no additional training, and the entire compression and inference process is performed on the GPU. Systematic experiments across diverse long-context tasks and compression ratios show that GraceKV ranks first in 24 of 32 settings and remains robust up to 128-fold compression. These results validate the effectiveness of global budget allocation in coordinating information coverage and local resolution.

↑