arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.32999cs.DC

Tessera:面向检索增强LLM服务的需求驱动KV缓存管理

Tessera: Demand-Driven KV Cache Management for Retrieval-Augmented LLM Serving

Fei Fang, Chung-Hsiang Lo, Yi Liu, Yifan Hua, Chen Qian

首次发表
浏览论文内容

中文总结 AI 辅助

Tessera通过将检索作为控制平面,实现需求驱动的KV缓存管理,解决RAG和智能体记忆中重复内容的KV重用问题,显著降低TTFT并保持答案质量。

中文摘要 AI 辅助

RAG和基于检索的智能体记忆都将检索到的内容注入LLM提示中,分别作为文档块和回忆的记忆记录。相同的内容可以在不同请求中以不同的提示位置或在不同的先前上下文之后重复出现,从而阻止通过传统前缀缓存进行重用。我们的特征分析发现,在智能体记忆工作负载中,在匹配前缀之外重复出现的记录占注入记忆令牌的70%以上。可组合的KV重用方法在这种情况下能够实现重用,但在线服务引入了一个管理问题:重复出现的单元其KV状态可能尚不存在、可能已被驱逐,或者可能驻留在另一个节点上。我们提出了Tessera,一个解耦的服务系统,使检索成为KV重用的控制平面。通过在模型执行前暴露所需的上下文单元,检索使Tessera能够将当前需求与检索历史、KV驻留状态和生成负载相结合,以协调缓存管理和请求路由。生成节点并发准备本地缓存、远程缓存和缺失的状态,同时将新计算的状态保留在请求的关键路径之外。在RAG和智能体记忆工作负载中,在匹配的请求速率下,Tessera相比使用EPIC的SGLang和LMCache,将平均TTFT降低了最多3.6倍,并在基线饱和的速率下维持低TTFT,同时匹配底层组合策略的答案质量。

英文摘要

RAG and retrieval-based agent memory both inject retrieved content into LLM prompts, as document chunks and recalled memory records, respectively. The same content can recur across requests at different prompt positions or after different preceding contexts, preventing reuse through conventional prefix caching. Our characterization finds that records recurring outside the matching prefix account for over 70% of injected memory tokens in agent-memory workloads. Composable KV-reuse methods enable reuse in such cases, but online serving introduces a management problem: a recurring unit's KV states may not yet exist, may have been evicted, or may reside on another node. We present Tessera, a disaggregated serving system that makes retrieval the control plane for KV reuse. By exposing the context units needed before model execution, retrieval allows Tessera to combine current demand with retrieval history, KV residency, and generation load to coordinate cache management and request routing. Generation nodes concurrently prepare locally cached, remotely cached, and missing states, while retaining newly computed states off the request's critical path. Across RAG and agent-memory workloads, Tessera lowers mean TTFT by up to 3.6x over SGLang and LMCache with EPIC at matched request rates, and sustains low TTFT at rates where the baselines saturate, while matching the answer quality of the underlying composition policy.

发表机构

  • University of California, Santa Cruz(加州大学圣克鲁兹分校)

机构由 AI 辅助整理,请以论文原文为准。

↑