arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

CacheReforge:适配器演化下陈旧KV缓存的有界恢复

CacheReforge: Bounded Recovery for Stale KV Caches under Evolving Adapters

Yuhang Cao, Yanzhou Mu, Chunrong Fang, Zhenyu Chen

arXiv 2609.30884首次发表:更新:

发表机构

Nanjing University(南京大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

CacheReforge通过逐层混合版本对象和累积尾部影响,在适配器演化下以最小有界重算恢复陈旧KV缓存,降低92.4%的KL散度并保留大部分缓存收益。

AI 中文摘要

大型语言模型依赖KV缓存来减少长上下文和交互式应用中的重复预填充计算。随着轻量级适配器的演化,缓存状态反映的是早期版本,因此陈旧缓存的重用会扭曲当前模型的输出,而完全重新计算受影响的后续部分虽能恢复保真度,但代价高昂。我们寻求最小化的重新计算来恢复当前适配器的行为。现有系统跟踪令牌、上下文或稳定的适配器身份,但既不能表示来自早期适配器版本的缓存,也不能区分更新传播与行为恢复所需的重新计算。为解决这些不足,我们引入CacheReforge,它将陈旧的KV缓存表示为逐层混合版本对象。它结合了逐层适配器锚点、校准的敏感性、累积漂移和可执行的重新启动边界,以选择直接重用、有界重新计算或完整的受影响后缀恢复。我们区分依赖深度与功能性重新计算范围,并使用累积尾部影响来刻画有界恢复何时能保持当前模型的行为。我们在Qwen2.5-1.5B和Qwen2.5-7B上使用持续的LoRA更新进行评估,包括16K HotpotQA和2WikiMQA工作负载。CacheReforge相对于陈旧重用将平均KL散度降低了92.4%,同时仅重新计算了5.44%的层,并将缓存维护时间相对于全新完整预填充减少了93.2%。这些结果表明,版本感知的恢复能保持模型保真度和大部分KV缓存的收益。

英文摘要

Large language models rely on KV caching to reduce repeated prefill computation in long context and interactive applications. As lightweight adapters evolve, cached states reflect earlier versions, so stale reuse distorts current model outputs, while complete affected suffix recomputation restores fidelity at substantial cost. We seek minimal recomputation that recovers current adapter behavior. Existing systems track token, context, or stable adapter identity, but neither represent caches from earlier adapter versions nor distinguish update propagation from the recomputation required for behavioral recovery. To address these gaps, we introduce CacheReforge, which represents stale KV caches as layerwise mixed-version objects. It combines per-layer adapter anchors, calibrated sensitivity, accumulated drift, and executable restart boundaries to select direct reuse, bounded recomputation, or complete affected-suffix recovery. We distinguish dependency depth from the functional recomputation horizon and use cumulative tail influence to characterize when bounded recovery preserves current-model behavior. We evaluate CacheReforge on Qwen2.5-1.5B and Qwen2.5-7B with continual LoRA updates, including 16K HotpotQA and 2WikiMQA workloads. CacheReforge reduces mean KL divergence by 92.4% relative to stale reuse, while recomputing only 5.44% of layers and reducing cache-maintenance time by 93.2% relative to fresh full prefill. These results show that version-aware recovery preserves model fidelity and most KV caching gains.

Comments13 pages, 5 figures. Artifact: https://doi.org/10.5281/zenodo.22951118

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑