arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.02901cs.LGcs.CL

AnchorKV:锚点-残差键值缓存压缩

AnchorKV: Anchor-Residual KV Cache Compression

Malik Khalaf, Yara Shamshoum, Nitzan Hodos, Yuval Sieradzki, Assaf Schuster

AI总结:

针对长上下文LLM推理的KV缓存内存瓶颈,提出AnchorKV压缩方案,不丢令牌且压缩20倍,70B规模下保留99%全缓存精度,大幅降低上下文成本。

AI中文摘要:

键值(KV)缓存是长上下文大语言模型(LLM)推理的主要内存瓶颈。现有方法从两端入手:驱逐方法永久丢弃令牌,若被丢弃的令牌后续被证明至关重要,就会降低性能;量化方法以低精度保留所有令牌,但压缩效果有限。我们提出AnchorKV,一种压缩方案,可将缓存压缩20倍,且不丢弃任何单个令牌。AnchorKV使用一组精确存储的小型锚点表示缓存,通过每个令牌与其最相似的锚点表达其他所有令牌,仅优化那些近似对模型输出影响最大的令牌。AnchorKV在模型和数据集上始终保持精度,在70B规模下保留了全缓存得分的99%,同时将整个上下文的成本降至其一小部分。

英文摘要:

The key-value (KV) cache is the primary memory bottleneck in long-context LLM inference. Existing approaches attack it from opposite ends: eviction methods permanently discard tokens, degrading performance whenever a discarded token later proves essential, while quantization methods retain all tokens at low precision but offer limited compression. We propose AnchorKV, a compression scheme that shrinks the cache by $20\times$ without discarding a single token. AnchorKV represents the cache using a small set of anchors stored exactly, expresses every other token through its most similar anchor, and refines only those whose approximation most affects the model's output. AnchorKV consistently preserves accuracy across models and datasets, retaining 99% of the full-cache score at the 70B scale, while keeping the entire context at a fraction of its cost.

补充信息

↑