KVCMAS:多智能体系统中共享上下文的高效KV缓存修正
KVCMAS: Efficient KV cache Correction for Shared Context in Multi-Agent Systems
查看机构详情
- Seoul National University(首尔大学)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
KVCMAS提出在线KV缓存修正框架,用低秩状态链式修正跨智能体缓存偏差,无需额外参考预填充,在保持准确性的同时实现2.0倍TTFT加速和最高3.7倍内存节省。
中文摘要 AI 辅助
面向提示词特化的多智能体系统允许多个智能体共享一个模型,同时执行互补角色以解决复杂任务。然而,智能体特定的前缀会改变为相同共享上下文生成的KV缓存,导致每个智能体重复预填充不断增长的上下文,并构建单独的缓存,带来高昂的计算和内存开销。选择性重计算减少了这种冗余,但仍保留了大量的模型执行,而现有的增量修正方法要么仅支持重复出现的上下文关系,要么为动态变化的上下文维护内存密集型的在线修正状态。对于首次出现的共享上下文,这些方法还会在智能体工作流之外构建参考缓存,并且第一个智能体处的近似修正会影响传递给后续智能体的输出。我们提出了KVCMAS,一种在线KV缓存修正框架,使用紧凑的低秩状态表示跨智能体的缓存偏差,并沿着智能体工作流无缝链式修正,无需额外的参考预填充。这种设计支持动态变化的共享上下文,同时保持第一个智能体缓存的精确性。在多种语言和视觉-语言工作负载上,KVCMAS在匹配或提高先前KV缓存共享方法准确性的同时,在高并发服务下实现了最低的TTFT。在受控服务轨迹下,与不共享KV缓存的推理相比,它提供了2.0倍的TTFT加速,并且相对于先前的KV缓存修正方法,峰值GPU内存最多减少了3.7倍。这些结果确立了KVCMAS作为一种面向提示词特化多智能体服务的准确且可扩展的KV缓存共享方法。
英文摘要
Prompt-specialized multi-agent systems enable multiple agents to share a model while performing complementary roles to solve complex tasks. However, agent-specific prefixes change the KV cache generated for the same shared context, causing each agent to repeatedly prefill the growing context and construct a separate cache with high computation and memory overhead. Selective recomputation reduces this redundancy but still retains substantial model execution, while existing delta correction methods either support only recurring context relations or maintain memory-intensive online correction states for dynamically changing context. For first seen shared context, these methods also construct a reference cache outside the agent workflow, and an approximate correction at the first agent affects the outputs passed to subsequent agents. We present KVCMAS, an online KV cache correction framework that represents cross-agent cache deviations using compact low-rank states and seamlessly chains corrections along the agent workflow without an additional reference prefill. This design supports dynamically changing shared context while preserving an exact first-agent cache. Across multiple language and vision-language workloads, KVCMAS matches or improves the accuracy of prior KV cache sharing methods while achieving the lowest TTFT under highly concurrent serving. Under controlled serving traces, it provides a 2.0x TTFT speedup over inference without KV cache sharing and reduces peak GPU memory by up to 3.7x relative to a prior KV cache correction method. These results establish KVCMAS as an accurate and scalable KV cache sharing approach for prompt-specialized multi-agent serving.