发表机构
Harvard University(哈佛大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究通过生产追踪和14种算法评估,发现前缀缓存中LRU已足够有效,并提出以最近性为基础、结合快速降级与计算感知淘汰的改进策略。
AI 中文摘要
长时间运行的LLM应用会反复发送不断增长的上下文,这使得前缀缓存对于降低预填充成本至关重要。然而,前缀缓存在智能体工作负载下的行为仍然知之甚少。我们研究了来自两家公司的生产环境追踪数据,并在HBM受限和大型内存池设置下评估了14种淘汰算法。尽管与Belady最优策略相比仍有较大差距,但为传统缓存设计的复杂策略相比LRU几乎没有带来额外收益。其原因是结构性的:前缀复用主要由活跃会话的规律性节奏主导,这使得最近性(recency)异常具有预测性。然而,前缀缓存也引入了新的挑战,包括重尾的会话足迹以及随着注意力计算随序列长度增长而高度可变的未命中成本。我们引入了计算节省比率和两个离线最优基准(oracle)来量化这些效应。结果表明,有效的前缀缓存管理应以最近性为基础,同时有选择地添加针对一次性前缀的快速降级、针对昂贵未命中的计算感知部分淘汰,以及依赖容量的淘汰粒度。我们将发布追踪数据和模拟器以支持未来研究。
英文摘要
Long-running LLM applications repeatedly send growing context, making prefix caching critical for reducing prefill cost. Yet prefix-cache behavior under agentic workloads remains poorly understood. We study production traces from two companies and evaluate 14 eviction algorithms across HBM-constrained and large memory-pool settings. Despite a large gap to Belady, sophisticated policies designed for traditional caches provide little benefit over LRU. The reason is structural: prefix reuse is dominated by the regular pacing of active sessions, making recency unusually predictive. Prefix caching nevertheless introduces new challenges, including heavy-tailed session footprints and highly variable miss costs as attention computation grows with sequence length. We introduce the compute-savings ratio and two offline oracles to quantify these effects. Our results show that effective prefix-cache management should retain recency as its foundation while selectively adding quick demotion for one-hit prefixes, compute-aware partial eviction for expensive misses, and capacity-dependent eviction granularity. We will release the traces and simulator to support future research.
Comments20 pages, 20 figures, 6 tables