arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.16215cs.AI

KV 缓存应存放于何处?面向长生命周期会话的 GPU、CPU 与 SSD 放置策略

Where Should the KV Cache Live? Placement Policies Across GPU, CPU, and SSD for Long-Lived Sessions

  • Vizuara

机构由 AI 辅助整理,请以论文原文为准。

Srikanta Datta Tumkur, Jay Iyer, Mehar Simhadri, Sai Pavan Kumar, Sai Kapil Kumar, Ramesh Nampelly

AI总结:

本研究通过离散事件模拟器评估KV缓存分层放置策略,发现分层容量(1+8+64)带来73倍并发提升,但预测重用与预取策略无效,最近使用与重用频率策略更优。

AI中文摘要:

GPU 高带宽内存稀缺且昂贵,而随着聊天、智能体循环和文档问答不断累积状态,KV 缓存会消耗其中大量内存。Mooncake、LMCache、FlexGen、InfiniGen 和 AttentionStore 等系统通过 CPU DRAM 和 SSD 扩展 GPU 内存。更难的问题在于哪些块应属于哪一层级、何时移动或驱逐它们,以及预取是否有帮助。我们在一个涵盖 GPU HBM、CPU DRAM 和 SSD 的离散事件模拟器中研究这些选择,并使用随机森林执行时间预测器进行校准。我们比较了聊天、智能体和文档问答工作负载下的最近使用(recency)、重用频率(reuse frequency)、预测重用(predicted reuse)以及带预取前瞻的 EWMA 预测器。分层支持每 GPU 并发会话数提升 73.02 倍,并将每会话成本降低 62.04 倍。这些收益来自 1 加 8 加 64 的分层容量,而非放置策略。在我们的设置中,批大小为 1 时解码受计算限制,因此放置对吞吐量影响甚微。它主要改变 PCIe 迁移流量和首 token 时间。对于聊天,最近使用产生的迁移流量比重用频率少 2.30 倍。重用频率在智能体和文档问答中表现最佳。现有的预测重用策略与最近使用在字节级别上相同,因此其智能体建议实际上等同于最近使用。真正的 EWMA 预测器改变了行为,但在预期其有帮助的工作负载上仍落后于重用频率。预取并不值得其带宽成本。在策略和缓存大小网格上,即使能预知未来请求的 oracle 在迁移流量上也从未优于不预取。特定于工作负载的放置可以减少数据移动,但预测重用和预取建议在实现中并未得到支持。

英文摘要:

GPU high bandwidth memory is scarce and expensive, and KV caches consume much of it as chats, agent loops, and document question answering accumulate state. Systems such as Mooncake, LMCache, FlexGen, InfiniGen, and AttentionStore extend GPU memory with CPU DRAM and SSD. The harder question is which blocks belong in each tier, when to move or evict them, and whether prefetching helps. We study these choices in a discrete event simulator spanning GPU HBM, CPU DRAM, and SSD, calibrated against a random forest execution time predictor. We compare recency, reuse frequency, predicted reuse, and an EWMA predictor with prefetch lookahead across chat, agent, and document question answering workloads. Tiering supports 73.02 times more concurrent sessions per GPU and lowers cost per session by 62.04 times. These gains come from tier capacities of 1 plus 8 plus 64, not placement policy. Decode is compute bound at batch size one in our setup, so placement barely affects throughput. It mainly changes PCIe migration traffic and time to first token. Recency produces 2.30 times less migration traffic than reuse frequency for chat. Reuse frequency performs best for agents and document question answering. The existing predicted reuse policy is byte identical to recency, making its agent recommendation effectively recency. A genuine EWMA predictor changes behavior but still ranks behind reuse frequency on the workloads prediction was expected to help. Prefetching does not justify its bandwidth cost. Across the policy and cache size grid, even an oracle with knowledge of future requests never beats no prefetch on migration traffic. Workload specific placement can reduce data movement, but the predicted reuse and prefetch recommendations are not supported as implemented.

↑