PatchKV:KV缓存的权重空间补偿
PatchKV: Weight-Space Compensation of KV Cache
查看机构详情
- KAIST(韩国科学技术院)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
PatchKV通过权重空间补偿,无需训练即可提升KV缓存压缩方法在高压缩比下的长上下文推理性能。
中文摘要 AI 辅助
大语言模型(LLM)的长上下文推理受限于键值(KV)缓存线性增长的内存需求。现有的压缩方法通过令牌驱逐或近似来减少缓存,但在高压缩比下性能急剧下降。我们提出PatchKV,一种无需训练的框架,通过将部分上下文携带到模型权重中来补偿KV缓存压缩方法。PatchKV将现成的压缩KV缓存与上下文特定的权重补丁配对,该补丁在上下文加载时计算一次,并用于该上下文的后续查询。权重补丁通过岭回归以闭式形式导出,通过对齐在完整缓存和压缩缓存下由上下文派生的参考查询令牌的块级激活。合并到模型后,该补丁在单上下文、多查询设置中不改变前向图和每次查询的推理成本。在三个模型架构上的长上下文问答(SCBench,最高170K令牌,SQuAD,NIAH)和数学(GSM8K)基准测试中,PatchKV持续改进缓存压缩方法,表明在高压缩比下补偿它们的一种替代方向。
英文摘要
Long-context inference with Large Language Models (LLMs) is bottlenecked by the linearly growing memory of the key-value (KV) cache. Existing compression methods reduce the cache through token eviction or approximation, but degrade sharply at aggressive compression budgets. We propose PatchKV, a training-free framework that compensates KV cache compression methods by carrying part of the context in the model's weights. PatchKV pairs an off-the-shelf compressed KV cache with a context-specific weight patch, which is computed once at context-loading time and served for downstream queries for the context. The weight patch is derived in closed form via ridge regression, by aligning the block-wise activations of context-derived reference query tokens under the full cache and the compressed cache. Once merged into the model, the patch leaves the forward graph and per-query inference cost unchanged in the single-context, multi-query setting. Across long-context QA (SCBench with up to 170K tokens, SQuAD, NIAH) and math (GSM8K) benchmarks on three model architectures, PatchKV consistently improves cache compression methods, suggesting an alternative direction to compensate them at aggressive budgets.