arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

ResKV:为固定预算KV缓存压缩重构被忽略的注意力贡献

ResKV: Reconstructing Omitted Attention Contributions for Fixed-Budget KV Cache Compression

Yuhang Zhan, Lisi Chen, Shuo Shang

arXiv 2607.29591首次发表:更新:

AI 中文总结

ResKV将固定KV预算分为主缓存和残差缓存,通过残差条目恢复注意力分子与分母质量,在LongBench等数据集上同预算下性能优于基线,兼顾压缩解码效率

AI 中文摘要

KV缓存压缩对于高效的长上下文推理至关重要。现有的驱逐方法会永久丢弃未被选中的token,从而消除它们对注意力的聚合贡献;而基于合并的替代方法虽能保留更多信息,但可能会扰动应保持精确的已保留键(Key)和值(Value)。我们发现,缓存驱逐所忽略的信息可被表述为softmax注意力分子和分母中的残差统计量。基于这一观察,我们提出了ResKV,它将固定的KV预算划分为精确的主缓存和紧凑的残差缓存,用于重构被忽略token的贡献。ResKV让主缓存token和残差条目参与相同的softmax归一化,因此残差条目可同时恢复注意力分子和分母的质量,而非作为事后校正。构建时的验证代理会为每一层和每个KV头确定残差分配,而解码时的动态门则会为单个查询调整残差贡献。在LongBench和RULER上进行的综合评估,涵盖了查询感知和查询无关设置、多种主干模型、缓存预算以及代表性压缩基线,结果表明,在相同的保留KV预算下,ResKV实现了广泛的性能提升,同时保持了压缩解码的实际效率,包括峰值内存使用量和长上下文解码吞吐量。

英文摘要

KV cache compression is essential for efficient long-context inference. Existing eviction methods permanently discard unselected tokens and consequently remove their aggregate contribution to attention. Merging-based alternatives preserve more information but can perturb retained keys and values that should remain exact. We observe that the information omitted by cache eviction can be formulated as residual statistics in both the numerator and denominator of softmax attention. Based on this observation, we propose ResKV, which divides a fixed KV budget into an exact main cache and a compact residual cache that reconstructs the contribution of omitted tokens. ResKV lets main-cache tokens and residual entries participate in the same softmax normalization, so residual entries restore both attention numerator and denominator mass rather than acting as a post-hoc correction. A construction-time validation proxy determines residual allocation for each layer and KV head, while a decode-time dynamic gate adjusts residual contributions for individual queries. Comprehensive evaluations on LongBench and RULER, covering query-aware and query-agnostic settings, multiple backbones, cache budgets, and representative compression baselines, demonstrate broad improvements under the same retained KV budget while preserving the practical efficiency of compressed decoding, including peak memory usage and long-context decode throughput.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑