发表机构
University of California, Santa Cruz; University of Washington; University of Chicago(加州大学圣克鲁兹分校; 华盛顿大学; 芝加哥大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对多智能体LLM服务中KV缓存复用不足的问题,提出CacheScout,通过在线学习智能体执行转换优化缓存管理,显著提升缓存命中率、降低延迟并提高吞吐量。
AI 中文摘要
多智能体大语言模型(LLM)系统已成为AI服务的重要部署范式,每个用户请求会被分解为一系列专用智能体。在这些工作流中,每个智能体重复执行由系统提示词、工具定义和少样本示例组成的固定上下文,这为KV缓存复用创造了大量机会。然而,现有的LLM服务系统采用前缀缓存和基于近期性的替换策略被动管理KV缓存,导致可复用的智能体上下文在下次调用前被驱逐,迫使重复计算。本文提出CacheScout,一种面向多智能体LLM服务的智能体感知KV缓存运行时层。核心洞见在于,未来KV缓存的复用仅由智能体执行语义而非缓存近期性决定。CacheScout通过在线学习智能体执行转换来捕获这些语义,无需预定义工作流图或离线训练,利用学习到的执行模型指导缓存驱逐和主动预取,同时不改变服务关键路径。我们在vLLM之上实现了CacheScout,在代表性真实多智能体工作负载上,CacheScout将KV缓存命中率提升10-18个百分点,降低平均首包时间(TTFT)18-45%,降低平均每轮延迟29-38%,并将峰值吞吐量提升最高57%。这些优势还可推广到更大模型,在维持37%更高吞吐量的同时,将TTFT降低最高54%。
英文摘要
Multi-agent LLM systems have emerged as an important deployment paradigm for AI services, where each user request is decomposed into a sequence of specialized agents. Across these workflows, every agent repeatedly executes a fixed context consisting of system prompts, tool definitions, and few-shot examples, creating substantial opportunities for KV-cache reuse. Existing LLM serving systems, however, manage KV-cache reactively using prefix caching and recency-based replacement, causing reusable agent contexts to be evicted before their next invocation and forcing repeated recomputation. We present CacheScout, an agent-aware KV-cache runtime layer for multi-agent LLM serving. The key insight is that future KV-cache reuse is governed by agent execution semantics rather than cache recency alone. CacheScout captures these semantics by learning agent execution transitions online, without requiring predefined workflow graphs or offline training, and uses the learned execution model to guide both cache eviction and proactive prefetching while leaving the serving critical path unchanged. We implement CacheScout on top of vLLM. Across representative real-world multi-agent workloads, CacheScout improves KV-cache hit rate by 10-18 percentage points, reduces mean TTFT by 18-45%, lowers mean per-turn latency by 29-38%, and increases peak throughput by up to 57%. These benefits also generalize to larger models, reducing TTFT by up to 54% while sustaining 37% higher throughput.