AI 中文总结
RestoreKV通过学习型恢复补充与查询无关的KV缓存驱逐方案,仅优化0.4%参数,在多主干、多基准上大幅降低压缩性能下降,提升长上下文处理能力且开销极小。
AI 中文摘要
与查询无关的KV缓存驱逐会对上下文进行一次压缩,并将生成的缓存重用于任意未来查询,但在严格预算下性能可能崩溃。现有方法主要改进保留哪些原始KV对。我们引入RestoreKV,在相同的总KV预算下,通过学习型恢复来补充这种基于选择的方案。我们的关键见解是,尽管驱逐造成的信息丢失是上下文特定的,但生成其紧凑补充的机制可跨上下文共享。上下文预填充后,少数恢复token通过一次经LoRA适配的传递关注完整KV缓存,生成紧凑的、上下文条件化的恢复缓存。基础重要性评分器和驱逐规则保持不变,且所有后续查询和解码时禁用适配器。RestoreKV通过来自冻结完整缓存模型的参数高效自蒸馏进行训练,仅优化0.4%的参数,且无需特定任务调优。在四个主干模型和四个长上下文基准上,RestoreKV大幅降低了压缩导致的性能下降。在Qwen3-4B上,它在五种基础驱逐方法的60对预算匹配设置中改善了59种;在5%预算下,它在RULER-4K上将KVzip从38.2提升至73.2;应用于KVzip+时,它在KVPress基准的16倍压缩下达到86.4的RULER准确率,同时在32K上下文评估中添加的一次性缓存构建开销小于0.5%。我们的项目页面可在此https URL获取。
英文摘要
Query-agnostic KV cache eviction compresses a context once and reuses the resulting cache for arbitrary future queries, but performance can collapse under tight budgets. Existing methods primarily improve which original KV pairs are retained. We introduce RestoreKV, which complements this selection-based formulation with learned restoration under the same total KV budget. Our key insight is that, although the information lost through eviction is context-specific, the mechanism for generating its compact complement can be shared across contexts. After context prefill, a few restore tokens attend to the full KV cache in a single LoRA-adapted pass, generating a compact, context-conditioned restore cache. The base importance scorer and eviction rule remain unchanged, and the adapters are disabled for all subsequent queries and decoding. RestoreKV is trained through parameter-efficient self-distillation from the frozen full-cache model, optimizing only $0.4\%$ of the parameters and requiring no task-specific tuning. Across four backbones and four long-context benchmarks, RestoreKV substantially reduces compression-induced degradation. On Qwen3-4B, it improves 59 of 60 paired, budget-matched settings across five base eviction methods; at a $5\%$ budget, it raises KVzip from $38.2$ to $73.2$ on RULER-4K. Applied to KVzip+, RestoreKV reaches $86.4$ RULER accuracy at $16\times$ compression on the KVPress Benchmark, while adding less than $0.5\%$ one-time cache-construction overhead in a 32K-context evaluation. Our project page is available at https://paper.pnu-cvsp.com/RestoreKV/
Comments13 pages, 8 figures