arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

PatchKV:动态编辑的LLM上下文的高效KV缓存恢复

PatchKV: Efficient KV Cache Recovery for Dynamically Edited LLM Contexts

Guotao Yang, Rui Guo, Siwei He, Sheng Chen, Yitao Hu, Keqiu Li

arXiv 2609.26219首次发表:更新:

发表机构

Tianjin University(天津大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

PatchKV通过配置文件引导的脏区域预测和块级精度恢复,实现动态编辑上下文中KV缓存的高效恢复,显著加速长上下文问答。

AI 中文摘要

长时间运行的LLM智能体工作流常常在保留长后缀的同时修改内部上下文片段。尽管后缀令牌保持不变,但改变的因果历史和旋转位置阻碍了对其卸载的键值(KV)状态的精确重用。完全重新计算后缀浪费了预填充工作,而不加区分地重用则传播过时状态,全精度恢复则增加了数据移动。我们提出了PatchKV,一种针对后缀保留修订的配置文件引导的恢复系统。PatchKV将相邻上下文版本分解为精确前缀、更新片段和对齐后缀。它使用离线长度条件漂移模型预测编辑局部脏区域,用从存储的注意力中选择的稀疏非局部块增强该区域,并将其并集块舍入为固定修复集。剩余的后缀块使用冻结的每块精度标签和用于反量化、RoPE校正和KV页放置的融合路径从CPU内存恢复。在三个模型和三个长上下文问答工作负载中,与完全后缀重新计算相比,PatchKV在平均恢复首令牌时间上实现了2.51-3.85倍的加速,与CacheBlend相比实现了1.26-2.06倍的加速,同时在九个设置中的六个中匹配或超过CacheBlend的F1分数,在其余设置中保持在1.36分以内。

英文摘要

Long-running LLM agent workflows often revise interior context spans while retaining long suffixes. Although suffix tokens remain unchanged, altered causal histories and rotary positions prevent exact reuse of their offloaded key-value (KV) states. Full suffix recomputation wastes prefill work, while indiscriminate reuse propagates stale states and full-precision restoration adds data movement. We present PatchKV, a profile-guided recovery system for suffix-preserving revisions. PatchKV decomposes adjacent context versions into an exact prefix, an updated span, and an aligned suffix. It predicts an edit-local dirty region using an offline length-conditioned drift model, augments this region with sparse nonlocal blocks selected from stored attention, and block-rounds their union into a fixed repair set. The remaining suffix blocks are restored from CPU memory using frozen per-block precision tags and a fused path for dequantization, RoPE correction, and KV-page placement. Across three models and three long-context question-answering workloads, PatchKV achieves a $2.51$-$3.85\times$ speedup in mean resume time-to-first-token over full suffix recomputation and a $1.26$-$2.06\times$ speedup over CacheBlend, while matching or exceeding CacheBlend's F1 score in six of nine settings and remaining within 1.36 points in the others.

Comments10 pages, 10 figures, 2 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑