arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

KV-Kaizen:学习上下文自适应的缓存压缩选择

KV-Kaizen: Learning Context-Adaptive Cache Compression Choices

Joao Monteiro, Louis Béthune, Anastasiia Filippova, Sonia Laguna, David Grangier, Marco Cuturi

arXiv 2609.37988首次发表:更新:

发表机构

Apple(苹果公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

KV-Kaizen通过学习上下文自适应的逐层缓存配置选择,在深度、精度和秩三个维度上组合压缩,无需驱逐令牌,在保持准确性的同时大幅减少解码缓存,支持大模型预训练后压缩。

AI 中文摘要

随着使用LLM处理的文本上下文规模增长,KV缓存的大小可能超过为原始模型权重分配的内存。这会对LLM的吞吐量产生负面影响,因为解码受内存限制,且解码成本随缓存大小增加而增长。近期工作通过丢弃最不相关的令牌来缓解这一瓶颈。驱逐引入了权衡,因为一次性丢弃内容的决定可能后来被证明是有害的。相反,我们专注于可以无需驱逐令牌即可实现缓存压缩的替代选择。我们通过学习一个选择器来实现这一点,该选择器能够基于上下文,为整体压缩预算生成每层缓存配置。选择器沿三个轴操作:跨层共享一个缓存(深度)、以更少比特缓存(精度)、或截断低秩潜在缓存表示(秩)。我们将所得方法称为KV-Kaizen,因为它复合了许多小的逐层选择。我们观察到,这些干预措施独立且均匀地应用于所有层时,会限制可实现的压缩,因为它们降低准确性。关键在于,局部且自适应于上下文的组合可以保持准确性,同时实现大量内存节省。在推理时,选择器在预填充之前运行一次。在指令遵循和推理任务的评估中,我们的选择器达到了准确性与缓存大小的帕累托前沿,优于无学习和事后基线。在长上下文任务上,KV-Kaizen改进了驱逐,并可与之组合,在14B模型上实现32倍更小的解码时缓存,同时保持准确性。4倍缓存大小缩减从7B参数起不产生准确性下降,且压缩模型比具有相同缓存大小的较小未压缩模型更准确。这些发现共同支持预训练大型模型,并仅在其后压缩它们。

英文摘要

As the context size of text processed with an LLM grows, the size of KV caches can outstrip the memory allocated for the original model weights. This impacts LLM throughput negatively, since decoding is memory-bound and decode cost grows with cache size. Recent work alleviates this bottleneck by discarding the least relevant tokens. Eviction introduces a tension, since a one-off decision to discard content may prove detrimental later. Instead, we focus on alternative choices that can lead to cache compression without evicting tokens. We achieve this by learning a selector that is able to produce, based on context, a per-layer cache configuration towards an overall compression budget. The selector operates along three axes: sharing one cache across layers (depth), caching at fewer bits (precision), or truncating the low-rank latent cache representations (rank). We call the resulting method KV-Kaizen, for the many small per-layer choices it compounds. We observe that these interventions taken independently and uniformly over all layers limit achievable compression because they degrade accuracy. Crucially, composing them locally and adaptively to the context can instead preserve accuracy while achieving large memory savings. At inference, the selector runs once, before pre-fill. In evaluations on instruction following and reasoning tasks, our selectors reach the Pareto frontier of accuracy against cache size, against learning-free and post-hoc baselines. On long-context tasks, KV-Kaizen improves on eviction and can be composed with it, reaching a 32x smaller decode-time cache on a 14B model while preserving accuracy. A 4x cache size reduction incurs no accuracy degradation from 7B parameters up, and a compressed model is more accurate than a smaller uncompressed one with the same cache size. Together, these findings support pre-training large models and compressing them only afterwards.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑