arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.32278cs.DC

RR-Evict:面向智能体LLM服务中超越LRU的细粒度前缀缓存逐出策略

RR-Evict: Fine-Grained Prefix Cache Eviction beyond LRU for Agentic LLM Serving

  • University of California San Diego(加州大学圣迭戈分校)

机构由 AI 辅助整理,请以论文原文为准。

Zaifeng Pan, Chris Wu, Zhengding Hu, Xinwei Qiang, Zhongkai Yu, Yufei Ding

AI总结:

针对智能体LLM服务中LRU逐出导致部分智能体上下文完全失效的问题,提出RR-EVICT细粒度前缀缓存逐出策略,通过轮询逐出各空闲智能体尾部块,保留更多可重用前缀,显著降低P99 TTFT和未缓存令牌。

AI中文摘要:

基于LLM的智能体通过重复的模型调用执行长时间跨度的任务,这些调用与工具执行和用户交互交替进行。由于每次调用都会扩展前几轮累积的历史记录,前缀缓存避免了智能体整个上下文的重复预填充。然而,聚合缓存占用空间随上下文长度和并发性增长,迫使服务系统回收缓存的KV张量。我们识别出“最近性同步”现象:对智能体缓存历史的访问会同时刷新其缓存节点,导致最近最少使用(LRU)逐出策略集中于少数智能体的私有历史。当被完全逐出的智能体返回时,它必须重新计算几乎整个累积上下文,即使大多数其他请求保留了大量的缓存重用,也会产生较大的首令牌时间(TTFT)异常值。我们提出了RR-EVICT,一种细粒度的前缀缓存逐出策略,将回收分散到空闲的智能体轨迹上。RR-EVICT以轮询方式访问智能体,并从每个智能体中逐出一个尾部块,当容量保留给空闲状态时,为更多智能体保留可重用的前缀。返回的智能体可以重用这些部分历史,并重新计算较小的缺失后缀。该策略无需预测未来的到达或工具延迟。我们在SGLang中实现了RR-EVICT,并在同地部署和预填充-解码分离部署下评估了对话和编码智能体工作负载。与LRU相比,RR-EVICT将P99 TTFT降低了高达75.4%,将P99未缓存提示令牌降低了高达65.7%。

英文摘要:

LLM-based agents execute long-horizon tasks through repeated model calls interleaved with tool execution and user interaction. As each call extends the history accumulated in previous turns, prefix caching avoids repeated prefill of the agent's entire context. However, the aggregate cache footprint grows with context length and concurrency, forcing serving systems to reclaim cached KV tensors. We identify recency synchronization: accesses to an agent's cached history refresh its cache nodes together, causing least recently used (LRU) eviction to concentrate on a few agents' private histories. When a fully evicted agent returns, it must recompute nearly its entire accumulated context, producing a large time-to-first-token (TTFT) outlier even if most other requests retain substantial cache reuse. We present RR-EVICT, a fine-grained prefix-cache eviction strategy that distributes reclamation across idle agent trajectories. RR-EVICT visits agents in round-robin order and evicts a tail chunk from each, preserving reusable prefixes for more agents when capacity remains for idle state. Returning agents can reuse these partial histories and recompute smaller missing suffixes. The policy requires no prediction of future arrivals or tool latency. We implement RR-EVICT in SGLang and evaluate conversational and coding-agent workloads under colocated and prefill-decode-disaggregated serving. Compared with LRU, RR-EVICT reduces P99 TTFT by up to 75.4% and P99 uncached prompt tokens by up to 65.7%.

↑