arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.06286cs.CL

DeferKV:重新思考一次性KV缓存压缩的驱逐时机

DeferKV: Rethinking Eviction Timing for One-Shot KV Cache Compression

Zhe Wang, Jiakai Li, Yujia Sun, Rongzheng Wang, Shuang Liang

首次发表
浏览论文内容

中文总结 AI 辅助

针对长上下文LLM中KV缓存压缩的驱逐时机问题,提出DeferKV方法,将驱逐决策推迟至首个解码步骤并融合两侧观察,无需额外训练,在多个基准上提升压缩性能并保持低延迟。

中文摘要 AI 辅助

长上下文大语言模型(LLM)在广泛的任务中展现了强大的能力,但不断增长的KV缓存引入了大量的内存和推理开销。现有的一次性KV缓存压缩方法通常在预填充后立即进行不可逆的驱逐,而此时尚未获得来自实际生成阶段的任何信号。我们的定量分析表明,实际生成阶段早期的查询提供的注意力信号与后续解码注意力更为一致,其中最大的单步增益出现在预填充-解码边界。基于这一观察,我们提出了DeferKV,它将驱逐决策从预填充结束推迟到第一个实际解码步骤,并在时间上结合提示侧和解码侧的观察,从而更好地使KV重要性估计与后续生成需求对齐。DeferKV不需要额外的训练、草稿模型或未来查询预测模块,使其简单且易于部署。在LongBench、RULER和Needle-in-a-Haystack上的实验表明,DeferKV在KV缓存压缩下持续提高模型性能,同时保持低推理延迟。

英文摘要

Long-context large language models (LLMs) have demonstrated strong capabilities across a wide range of tasks, but the growing KV cache introduces substantial memory and inference overhead. Existing one-shot KV cache compression methods typically commit to irreversible eviction immediately after prefill, before any signal from actual generation becomes available. Our quantitative analysis shows that early queries from the actual generation stage provide attention signals that are more consistent with subsequent decode attention, with the largest single-step gain occurring at the prefill-decode boundary. Based on this observation, we propose DeferKV, which moves the eviction decision from the end of prefill to the first real decoding step and temporally combines prompt-side and decode-side observations, thereby better aligning KV importance estimation with subsequent generation requirements. DeferKV requires no additional training, draft model, or future-query prediction module, making it simple and easy to deploy. Experiments on LongBench, RULER, and Needle-in-a-Haystack demonstrate that DeferKV consistently improves model performance under KV cache compression while maintaining low inference latency.

发表机构

  • University of Electronic Science and Technology of China(电子科技大学)
  • Ubiquitous Intelligence and Trusted Services Key Laboratory of Sichuan Province(四川省泛在智能与可信服务重点实验室)

机构由 AI 辅助整理,请以论文原文为准。

↑