邻接性而非重要性:文档编辑后陈旧KV缓存的有预算修复
Contiguity, Not Importance: Budgeted Repair of Stale KV Caches After Document Edits
- University of South Florida(南佛罗里达大学)
- DEVCOM Army Research Laboratory(DEVCOM陆军研究实验室)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文研究文档编辑后KV缓存修复问题,提出有预算的连续编辑局部窗口重算策略,在事实RAG基准上恢复高比例答案余量,比完全重预填充快13-21倍,并揭示邻接性而非重要性决定修复效果。
AI中文摘要:
KV缓存重用可以降低检索增强生成和智能体系统中的推理成本,但当检索到的知识、工作记忆或用户状态被编辑时,缓存的上下文可能变得过时。在因果自注意力机制下,即使是局部编辑也可能影响下游的KV状态。完全重新预填充能可靠地恢复一致性,但代价高昂,而仅刷新被编辑的跨度可能会使下游依赖关系过时。我们将原地修复形式化为有预算的重新计算,并在一个事实性RAG基准上比较了无训练的位置选择策略,该基准包含匹配的直接编辑和派生编辑。在三个模型家族中,所有策略都能修复直接情况,但派生情况明显将它们区分开来。在主要预算下,连续的编辑局部窗口恢复了至少0.94的编辑后答案余量,并显著优于基于注意力、KV偏差和结构的选择器。机制分析表明,在干净状态移植下有效的位置集合在实际重新计算中可能失败,因为分散的位置继承了周围的过时性。编辑局部的优势也依赖于邻接性,当承载答案的文本向下游移动时,这种优势基本消失。由于与答案相关的编辑几乎总是破坏模型行为,失败严重性难以预测,且修复比完全重新预填充快13-21倍,我们的结果支持在依赖文本保持与编辑相邻时进行无条件的编辑局部修复。
英文摘要:
KV-cache reuse can reduce inference cost in retrieval-augmented generation and agentic systems, but cached contexts may become stale when retrieved knowledge, working memory, or user state is edited. Under causal self-attention, even a local edit can affect downstream KV states. A full re-prefill reliably restores consistency but is costly, whereas refreshing only the edited span can leave downstream dependencies stale. We formulate in-place repair as budgeted recomputation and compare training-free position-selection policies on a factual RAG benchmark with matched direct and derived edits. Across three model families, all policies repair direct cases, but derived cases clearly separate them. At the primary budget, a contiguous edit-local window recovers at least 0.94 of the post-edit answer margin and substantially outperforms attention-based, KV-deviation, and structural selectors. Mechanistic analysis shows that position sets effective under clean-state transplantation can fail under actual recomputation because scattered positions inherit surrounding staleness. The edit-local advantage also depends on adjacency and largely disappears when the answer-bearing text moves downstream. Because answer-relevant edits almost always corrupt model behavior, failure severity is difficult to predict, and repair is 13-21 times faster than full re-prefill, our results support unconditional edit-local repair when the dependent text remains adjacent to the edit.