EfficientAgent:是什么让KV缓存卸载对并发智能体有效?
EfficientAgent: What Makes KV Cache Offloading Work for Concurrent Agents?
浏览论文内容
中文总结 AI 辅助
针对LLM智能体KV缓存卸载效果不一致的问题,提出EfficientAgent,通过栈距离模型估计重用工作集来调整主机层大小,并采用运行时策略管理写入,显著减少重新计算和端到端时间。
中文摘要 AI 辅助
大语言模型智能体在每一轮对话中都会重新发送整个对话历史,而其中大部分内容在上一轮已经处理过。服务系统通过缓存键值(KV)状态来避免重复计算,并在GPU内存不足时将该状态卸载到主机内存。对于智能体而言,卸载会产生不一致的结果:在相同的编码智能体工作负载上,它加速了一个部署,却减慢了另一个部署,并且在第三个部署上没有任何变化,即使将令牌加载回来的成本比重新计算它要便宜几倍。原因在于缓存状态必须存活到再次被使用。当一个智能体等待其工具时,服务器会处理所有其他智能体的上下文,因此一个智能体的前缀只有在主机层持有整个智能体池的可重用上下文(我们称之为重用工作集)时才会被重用。较小的层会不断写入状态,而这些状态在任何人读取之前就被驱逐了。我们提出了EfficientAgent,它通过这个工作集来调整和管理主机层的大小。一个栈距离模型根据智能体历史估计工作集以确定主机层的大小;其实验前做出的预测定位了卸载开始产生收益的容量。当层太小时,运行时策略会停止写入被驱逐上下文的大的重填,并继续扩展仍然被缓存的前缀;当层足够大时,它会写入所有内容。在SWE-bench Verified编码智能体上,根据估计的工作集调整大小的主机层将重新计算的提示令牌减少了93%,端到端时间减少了39%。对于小的固定层,该策略将重新计算减少了35%;对于大的层,它避免了因始终过滤写入而导致的4.3倍增加。在三种GPU类型和两种模型上,当GPU相对于主机带宽的字节计算能力较低且主机层持有工作集时,卸载是有益的。代码可在该https URL获取。
英文摘要
LLM agents resend their whole conversation on every turn, and most of it was already processed on the previous turn. Serving systems avoid recomputing it by caching its key-value (KV) state and, when GPU memory runs out, by offloading that state to host memory. For agents, offloading gives inconsistent results: on the same coding-agent workload it speeds up one deployment, slows down another, and changes nothing on a third, even where loading a token back is several times cheaper than recomputing it. The reason is that cached state must survive until it is used again. While one agent waits for its tool, the server processes the contexts of all other agents, so an agent's prefix is reused only if the host tier holds the reusable context of the whole agent pool, which we call the reuse working set. A smaller tier keeps writing state that is evicted before anyone reads it. We present EfficientAgent, which sizes and manages the host tier by this working set. A stack-distance model estimates the working set from agent histories to size the host tier; its predictions, made before the experiments, located the capacity at which offloading starts to pay. When the tier is too small, a runtime policy stops writing large refills of evicted context and keeps extending prefixes that are still cached; when the tier is large enough, it writes everything. On SWE-bench Verified coding agents, a host tier sized to the estimated working set cuts recomputed prompt tokens by 93% and end-to-end time by 39%. With a small fixed tier, the policy cuts recomputation by 35%; with a large tier, it avoids the 4.3-fold increase caused by always filtering writes. Across three GPU types and two models, offloading pays off when the GPU has little compute per byte of host bandwidth and the host tier holds the working set. Code is available at https://github.com/KunmingSHAO/efficientagent_release.
发表机构
- The Hong Kong University of Science and Technology(香港科技大学)
- Huawei Technologies Ltd.(华为技术有限公司)
- The University of Hong Kong(香港大学)
- Sun Yat-sen University(中山大学)
机构由 AI 辅助整理,请以论文原文为准。