arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

构建 py-kvcache:vLLM 在 NVMe SSD 上的外部 KV 缓存性能表征

Building py-kvcache: A Performance Characterization of External KV Caching for vLLM with NVMe SSDs

Joseph Kanichai, Tiziano De Matteis, Animesh Trivedi

arXiv 2609.11744首次发表:更新:

发表机构

Vrije Universiteit Amsterdam; IBM Research Zurich(阿姆斯特丹自由大学; IBM苏黎世研究院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究通过构建 py-kvcache 连接器,利用异步 I/O 和调度器感知预加载,在 NVMe SSD 上实现外部 KV 缓存,显著提升 vLLM 长上下文请求的 TTFT 性能,并强调缓存决策需依具体硬件而定。

AI 中文摘要

前缀缓存可以通过重用先前计算的关键-值(KV)状态来减少长上下文 LLM 请求的首 token 时间(TTFT),但对于短前缀或快速 GPU,重新计算可能比从外部缓存加载更快。我们在 vLLM 中通过合成工作负载、长上下文基准测试和生产轨迹,对 GPU、CPU 和 NVMe 三层之间的这种权衡进行了表征,发现缓存性能取决于传输粒度、中间内存使用以及传输进入请求调度的时间,而不仅仅取决于设备带宽。这些发现促使我们开发了 py-kvcache,一个 vLLM KV 卸载连接器,具有异步直接 I/O、有界共享暂存和调度器感知预加载,它在请求仍在等待时开始磁盘读取,与计算重叠。在 80k token 时,py-kvcache 从磁盘加载比 LMCache 快 2.0 倍,其中预加载贡献了 1.34 倍。启用 GPU、CPU 和磁盘缓存后,它比 LMCache 快 1.23 倍,并且与原生 vLLM KV 卸载实现相差约 4%。LongBench 和 SCBench 表明这些优势扩展到不规则前缀链和多轮工作负载。Bailian 轨迹重放在较弱的 GPU 上改善了 TTFT,但在 H100 上,平均请求低于盈亏平衡点,仅 GPU 内存就足以保留足够的前缀。因此,外部 KV 缓存应被视为特定设置的准入决策。py-kvcache 实现可在以下网址获取:此 https URL。

英文摘要

Prefix caching can reduce the time to first token (TTFT) of long-context LLM requests by reusing previously computed key-value (KV) states, but for short prefixes or fast GPUs, recomputation can be faster than loading from an external cache. We characterize this tradeoff in vLLM across GPU, CPU, and NVMe tiers using synthetic workloads, long-context benchmarks, production traces, and find that cache performance depends on transfer granularity, intermediate memory use, and when transfers enter the request schedule, not only on device bandwidth. These findings motivate py-kvcache, a vLLM KV Offload connector with asynchronous direct I/O, bounded shared staging, and scheduler-aware preloading, which starts disk reads while requests are still waiting, overlapping with compute. At 80k tokens, py-kvcache loading from disk is 2.0x faster than LMCache, with preloading contributing 1.34x. With GPU, CPU, and disk caching enabled, it is 1.23x faster than LMCache and within approximately 4% of the native vLLM KV Offload implementation. LongBench and SCBench show that these benefits extend to irregular prefix chains and multi-turn workloads. Bailian trace replays improve TTFT on a weaker GPU, but on an H100 the average request falls below the break-even point and GPU memory alone retains enough prefixes. External KV caching should therefore be treated as a setup specific admission decision. The py-kvcacheimplementation is available at: https://github.com/atlarge-research/py-kvcache.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑