发表机构
Carnegie Mellon University; Capital One; University of Chicago; University of California, Berkeley; University of Southern California; University of Washington(卡内基梅隆大学; 第一资本; 芝加哥大学; 加州大学伯克利分校; 南加州大学; 华盛顿大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究揭示滚动智能体场景下KV缓存重用存在历史依赖问题,提出文档对齐重计算策略,在5%预算下将答案变异从69.0%降至26.1%,保真度提升34.5-52.5个百分点,并保持约5.7倍加速。
AI 中文摘要
长期运行的智能体会反复调用大语言模型,同时保留其文档窗口的大部分内容,逐出旧文档并追加新文档。这些滚动更新破坏了精确前缀缓存,并促使采用带有选择性重计算的非前缀KV缓存重用。我们表明,带有选择性重计算的持久KV缓存重用可能是历史依赖的:在我们的滚动智能体工作负载中,一个未改变的提示词可能因之前处理的请求不同而产生不同的答案。在匹配的5%重计算预算下,文档对齐的重计算将跨请求顺序的答案变异从CacheBlend的token top-k策略下的69.0%降低到26.1%。当每个提示词在不同顺序的前置请求序列之后进行评估时,文档对齐的重计算相比token top-k策略,将相对于完整预填充的保真度提高了34.5-52.5个百分点,而两种策略均实现了约5.7倍的中位TTFT加速。我们的消融研究表明,在我们的滚动智能体工作负载中,连续性是与稳健选择性重计算相关的主要因素。
英文摘要
Long-running agents repeatedly call an LLM while retaining most of their document window, evicting old documents, and appending new ones. These rolling updates break exact prefix caching and motivate non-prefix KV-cache reuse with selective recomputation. We show that persistent KV-cache reuse with selective recomputation can be history-dependent: in our rolling-agent workload, an unchanged prompt can produce different answers depending on the requests processed before it. At a matched 5% recomputation budget, document-aligned recomputation reduces answer variation across request orders from 69.0% with CacheBlend's token top-$k$ policy to 26.1%. When each prompt is evaluated after a different sequence of preceding requests, document-aligned recomputation improves fidelity to full prefill by 34.5-52.5 percentage points over token top-$k$, while both policies achieve approximately 5.7$\times$ median TTFT speedup. Our ablation study shows that, in our rolling-agent workload, contiguity is the main factor associated with robust selective recomputation.
Comments10 pages, 5 figures