发表机构
Xiamen University; WeChat AI, Tencent Inc.(厦门大学; 腾讯微信人工智能实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
CommitKV通过提交转换识别KV生命周期,区分休眠与可安全移除信息,在多轮智能体中降低内存占用、加速推理且准确率优于现有KV缓存压缩方法。
AI 中文摘要
多轮推理与行动(ReAct)智能体积累的推理、工具调用及观测轨迹会不断增长,其键值(KV)缓存也随之增大,进而增加了模型推理时的内存占用与注意力成本。现有的KV缓存压缩方法通过移除注意力分数较低的状态来降低这些成本,但当前轮次的低注意力并不意味着未来无关,因为暂时不活跃的信息可能在后续变得重要,因此基于快照的移除方法未明确区分暂时休眠的信息与已完成作用的信息。本文提出CommitKV,其通过提交转换识别KV生命周期:首先将已完成的智能体事件划分为token页,对比每个符合条件的页面在工具调用提交前、以及该提交返回的观测值被纳入后删除的影响;基于这些配对测量结果,CommitKV将休眠页面与高完成度候选页面区分开,再应用贪心联合测试,仅当候选页面的联合提交后影响仍有限时,才接受其退役。最后,在后续压缩检查点,已接受的页面被排除,一组有限的等待提交后测量的页面被保护,其余KV状态则在缓存预算内保留,使用相同的token索引作为键、值和绝对位置。这些机制确保CommitKV能区分休眠信息与已完成观测作用、可安全移除的信息。在多个基准上的实验表明,CommitKV可降低智能体内存占用、加速端到端推理,且比现有KV缓存压缩方法实现更高的准确率。
英文摘要
Multi-turn Reasoning-and-Acting (ReAct) agents accumulate growing trajectories of reasoning, tool calls, and observations. Their key-value (KV) caches grow accordingly, increasing memory use and attention cost during model inference. Existing KV cache compression methods reduce these costs by evicting states with low attention scores. However, low attention in the current turn does not imply future irrelevance, as temporarily inactive information may become important later. Snapshot-based eviction methods therefore do not explicitly distinguish temporarily dormant information from information that appears to have completed its role. In this paper, we present CommitKV, which identifies KV lifecycles through commit transitions. Specifically, CommitKV first divides completed agent events into token pages and compares each eligible page's deletion effect before a tool-call commit and after the commit's returned observation has been incorporated. Based on these paired measurements, CommitKV distinguishes dormant pages from high-to-low completion candidates. It then applies a greedy joint test, accepting candidates for retirement only when their combined post-commit effect remains bounded. Finally, at a later compression checkpoint, accepted pages are excluded, a bounded set of pages awaiting post-commit measurement is protected, and the remaining KV states are retained within the cache budget using the same token indices for keys, values, and absolute positions. These mechanisms ensure that CommitKV can distinguish dormant information from information that has completed its observed role and can be safely removed. Experiments on various benchmarks show that CommitKV reduces agent memory use, accelerates end-to-end inference, and achieves higher accuracy than existing KV cache compression methods.