发表机构
Karlsruhe Institute of Technology(卡尔斯鲁厄理工学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
KV$^2$提出基于选择性重建的查询无关KV缓存压缩方法,通过轻量级代理评分器筛选重要标记后仅重处理子集,在低预算下显著提升长上下文模型性能,并降低运行时间和内存。
AI 中文摘要
键值(KV)缓存的内存占用限制了长上下文模型的实际应用,并且当一个预填充的上下文需要服务多个不同查询时,它主导了成本。在这种可重用场景中,查询无关的压缩在成本与质量之间进行权衡:轻量级估计器便宜但准确性较低,而全上下文重建评分更准确但会重新处理整个提示。我们引入了KV$^2$,一种基于选择性重建的查询无关KV缓存压缩方法。KV$^2$首先使用轻量级代理评分器识别信息丰富的上下文内标记,然后仅重新处理这一子集以计算最终驱逐分数。在RULER、Needle-in-a-Haystack和LongBench上,随着预算收紧,KV$^2$相对于基线的优势扩大:在RULER 16K上,在2%的KV缓存预算下,它比次优基线的平均得分提高了超过40个百分点;在LongBench上,在2%-10%的预算范围内,它获得了最高的平均得分,同时压缩阶段的运行时间和峰值内存低于全上下文重建。因此,可重用的KV缓存压缩不需要重新处理整个上下文。我们的代码可在以下https URL获取。
英文摘要
The memory footprint of the key-value (KV) cache constrains the practical use of long-context models, and it dominates cost when one prefilled context must later serve many different queries. In this reusable setting, query-agnostic compression trades cost against quality: lightweight estimators are cheap but less accurate, whereas full-context reconstruction scoring is more accurate yet reprocesses the entire prompt. We introduce KV$^2$, a query-agnostic KV-cache compression method based on selective reconstruction. KV$^2$ first uses a lightweight proxy scorer to identify informative in-context tokens, then reprocesses only this subset to compute final eviction scores. On RULER, Needle-in-a-Haystack, and LongBench, KV$^2$'s margin over baselines widens as the budget tightens: on RULER 16K at a 2% KV-cache budget it improves the average score over the next-best baseline by more than 40 percentage points, and on LongBench it attains the highest average across 2%-10% budgets at lower compression-stage runtime and peak memory than full-context reconstruction. Reusable KV-cache compression thus does not require reprocessing the full context. Our code is available at https://anonymous.4open.science/r/KVsquared-0B97.