arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

EchoPress:通过虚拟上下文重建实现查询无关的KV缓存剪枝

EchoPress: Query-Agnostic KV Cache Pruning via Virtual Context Reconstruction

Jiawei Lin, Saibo Geng, Thomas Bourgeat

arXiv 2610.00412首次发表:更新:

发表机构

EPFL(洛桑联邦理工学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

EchoPress提出无需训练的KV缓存剪枝方法,通过虚拟上下文重建近似注意力分数,在保持准确率的同时大幅降低压缩开销和预填充时间。

AI 中文摘要

KV缓存剪枝通过驱逐较不重要的键值对来减少长上下文推理的内存占用。KVzip通过上下文重建来估计重要性:提示模型逐块重复上下文。这以额外的前向传播为代价实现了强大的压缩质量。学习型近似方法降低了这一成本,但需要针对特定模型的训练。我们分析了KVzip如何识别重要的缓存信息,并展示了如何利用预填充期间已计算的信息来近似其重建分数。这些发现催生了EchoPress,一种无需训练的方法,它使用标准预填充中的查询和键来近似重建注意力。对于每个请求,它仅重建第一个块以校准剩余上下文的重要性分数。在LongBench和RULER上使用Qwen3-8B和Llama-3.1-8B-Instruct进行的实验表明,在50%至90%的驱逐比率下,EchoPress在任务准确率上与KVzip相当,同时将压缩开销降低了1.7-19.6倍,并将总预填充时间降低了最多2.9倍。代码可在该https URL获取。

英文摘要

KV cache pruning reduces long-context inference memory usage by evicting less important key-value pairs. KVzip estimates importance through context reconstruction: prompting a model to repeat the context chunk by chunk. This achieves strong compression quality at the cost of additional forward passes. Learned approximations reduce this cost but require model-specific training. We analyze how KVzip identifies important cached information and show how to approximate its reconstruction scores using information already computed during prefill. These findings motivate EchoPress, a training-free method that approximates reconstruction attention using queries and keys from standard prefill. For each request, it reconstructs only the first chunk to calibrate importance scores for the remaining context. Experiments on LongBench and RULER with Qwen3-8B and Llama-3.1-8B-Instruct show that EchoPress matches KVzip in task accuracy across eviction ratios from 50% to 90%, while reducing compression overhead by a factor of 1.7-19.6 and total prefill time by a factor of up to 2.9. Code is available at https://github.com/ljwljwljwljw/kvpress/tree/echo-press.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑