发表机构
University of California, San Diego(加利福尼亚大学圣地亚哥分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对内存增强语言模型智能体推理成本高的问题,提出无训练、探针引导的AgentKVShift方法,通过估计内存级偏移量校正重用令牌,在多语言模型和基准测试中表现出色,实现低重新计算量和预填充加速,还能与缓存量化正交组合提升性能。
AI 中文摘要
内存增强的语言模型智能体通过智能体记忆系统在数百次交互中维护上下文,该系统利用语言模型生成的元数据(如摘要、关键词和标签)积极管理检索到的内容。从推理成本的角度来看,每次检索都会触发这些结构化记忆单元完全重新编码为键值(KV)状态,这主导了预填充延迟。现有的无训练KV重用方法通过选择性地重新计算一小部分令牌来缓解此问题,但它们是为RAG风格的原始段落设计的,在结构化智能体记忆上会退化。我们提出了AgentKVShift,一种无训练、探针引导的KV残差校正方法,它针对每个检索到的记忆单元进行操作。我们证明的一个关键见解是,每个记忆的KV重用残差可分解为共享的内存级偏移量加上小的逐令牌波动。从小探针集中估计此偏移量使我们能够通过单个加权校正来校正每个重用的令牌。与之前的重用方法不同,AgentKVShift还会校正它不重新计算的令牌,将刷新预算转化为整个块中的有用信号。在跨越3B到\n32B参数的四个开源语言模型和两个长期智能体记忆基准(长期对话和智能体应用)上,AgentKVShift在仅刷新10%-30%的缓存时实现了接近完全重新计算的性能,在相同的重新计算比率下优于基线。它需要低5倍的重新计算量才能达到这种接近完全的性能,而之前的重用方法仅在4\n5%-55%的刷新时才能实现。在这种情况下,AgentKVShift在单个A100上比不进行KV重用的情况实现了2-3.5倍的预填充加速。AgentKVShift与KV缓存量化正交组合,在激进的2位和4位设置下保留了超过2倍于之前重用方法的F1值。
英文摘要
Memory-augmented LLM agents maintain context across hundreds of interactions through agentic memory systems that actively curate retrieved content with LLM-generated metadata such as summaries, keywords, and tags. From an inference cost standpoint, every retrieval triggers a full re-encoding of these structured memory units into Key-Value (KV) states, which dominates prefill latency. Existing training-free KV reuse methods mitigate this by selectively recomputing a small fraction of tokens, but were designed for RAG-style raw passages and degrade on structured agentic memories. We present AgentKVShift, a training-free, probe-guided KV residual correction method that operates per retrieved memory unit. One of the crucial insights we demonstrate is that the per-memory KV reuse residual decomposes into a shared memory-level offset plus small token-wise fluctuations. Estimating this offset from a small probe set allows us to correct every reused token by a single weighted correction. Unlike prior reuse methods which decide which tokens to recompute and leave the rest of the cache stale, AgentKVShift also corrects the tokens it does not recompute, turning the refresh budget into useful signal across the entire chunk. Across four open-source LLMs spanning 3B to 32B parameters and two long-horizon agentic memory benchmarks (long-term dialogue and agentic applications), AgentKVShift achieves near full recompute performance while refreshing only 10-30% of the cache, outperforming baselines at the same recompute ratio. It requires up to 5x lower recompute to reach this near-full performance, which prior reuse methods only attain at 45-55% refresh. In this regime, AgentKVShift delivers prefill speedups of 2-3.5x over no-KV-reuse on a single A100. AgentKVShift orthogonally composes with KV cache quantization, retaining over 2x the F1 of prior reuse methods under aggressive 2- and 4-bit settings.