arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.39819cs.OS

捕获生命周期:KVTether 在 ReAct 智能体中的 KV 缓存管理

Capture the lifecycle: KV Cache management in ReAct Agents with KVTether

Kaihua Fu, Yukun Zhou, Chaokun Chang, Yinghao Yu, Luping Wang, Guodong Yang, Jiuchen Shi, Quan Chen, Wei Wang

首次发表
浏览论文内容

中文总结 AI 辅助

KVTether 通过追踪智能体框架的语义原语,将消息级生命周期转化为 KV 级状态,实现状态优先的缓存管理,显著降低延迟与成本。

中文摘要 AI 辅助

高效服务长上下文推理与行动(ReAct)智能体依赖于 KV 缓存复用,以减少大语言模型(LLM)的预填充延迟和货币成本。然而,智能体框架与底层服务栈之间存在语义鸿沟。通过上下文变更、工具执行和子智能体协调,上下文消息可能变得活跃参与、被永久丢弃或暂时未使用,而服务栈仅观察到对应 KV 缓存的访问。这种生命周期盲区使得诸如 LRU 等仅基于最近使用(recency-only)的策略无法及时回收已死亡的 KV,也无法保留那些比新条目更早被复用的旧 KV。我们提出 KVTether,一个面向 ReAct 智能体的生命周期感知 KV 缓存管理框架。通过追踪嵌入在智能体框架中的语义原语,KVTether 在高度动态的执行过程中捕获运行时生命周期语义。随后,KVTether 将消息级语义转换为 KV 级生命周期状态,并利用这些状态驱动状态优先的缓存管理,而无需向智能体框架暴露物理复杂性。在回收已死亡的 KV 后,KVTether 优先保留等待复用的存活但空闲(live-but-idle)KV,减少在复用前的过早驱逐。在智能体基准测试和生产工作负载中,相对于 LMCache 和 MORI,KVTether 分别将端到端请求延迟降低高达 26.3% 和 17.4%,并将估计的任务成本平均降低 40.0% 和 33.2%。

英文摘要

Efficient serving of long-context reasoning-and-acting (ReAct) agents relies on KV cache reuse to reduce large language model (LLM) prefill latency and monetary cost. However, a semantic gap exists between agent harnesses and the underlying serving stack. Through context mutation, tool execution, and subagent coordination, context messages may become actively engaged, permanently discarded, and temporarily unused, while the serving stack only observes accesses to the corresponding KV cache. This lifecycle blindness prevents recency-only policies such as LRU from reclaiming dead KV promptly and from preserving older KV that will be reused sooner than newer entries. We present KVTether, a lifecycle-aware KV cache management framework for ReAct agents. By tracing semantic primitives embedded in agent harnesses, KVTether captures runtime lifecycle semantics during highly dynamic execution. KVTether then translates message-level semantics into KV-level lifecycle states and uses these states to drive state-prioritized cache management without exposing physical complexities to agent harnesses. After reclaiming dead KV, KVTether preferentially preserves live-but-idle KV that is waiting for reuse, reducing premature eviction before reuse. Across agent benchmarks and production workloads, KVTether reduces end-to-end request latency by up to 26.3% and 17.4% relative to LMCache and MORI, respectively, and lowers estimated task cost by 40.0% and 33.2% on average.

↑