TempoKV:面向内存语义闪存的LLM KV缓存及时暂存
TempoKV: Timely Staging of LLM KV Caches for Memory-Semantic Flash
- Samsung Semiconductor(三星半导体)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
TempoKV通过时序感知的资源承诺层,将KV缓存复用认知与暂存资源获取分离,在SSD内存语义闪存上减少快速层占用,同时保持服务性能。
AI中文摘要:
在大型语言模型(LLM)服务中,可复用的前缀键值(KV)缓存可能超出GPU内存容量。内存语义闪存层次结构提供了基于SSD的容量,但快速层容量有限,然而逻辑上的KV命中并不一定已准备好供GPU检索。按需暂存会暴露SSD延迟,而立即暂存则可能在检索开始前很久就预留快速层容量。我们提出了TempoKV,一个时序感知的资源承诺层,它将复用的早期认知与暂存资源的获取分离。它仅以元数据声明的方式记录可复用KV命中,并在运行时估计的检索时间降至存储估计的使KV驻留并防止驱逐所需时间时请求承诺。这些估计适应运行时进度和暂存状态,而承诺仍受可用受保护容量的约束。我们在基于SSD的CXL内存设备上,在vLLM和LMCache中实现了TempoKV,且不改变请求调度。在两种模型和三种前缀缓存比率下,与立即暂存相比,TempoKV将每个请求的受保护快速层字节时间减少了63-91%,同时保留了提前暂存的大部分服务优势。在快速层容量扫描中,当容量从100 GiB降至25 GiB时,输出吞吐量和p95首令牌时间(TTFT)几乎保持不变。与未修改的LMCache的Device-DAX L1配置相比,TempoKV将p95 TTFT最多降低48.0%,并将输出吞吐量最多提高27.8%。
英文摘要:
Reusable prefix key-value (KV) caches can outgrow GPU memory in large language model (LLM) serving. A memory-semantic flash hierarchy offers SSD-backed capacity with a limited fast tier, but a logical KV hit is not necessarily ready for GPU retrieval. Demand staging exposes SSD latency, whereas immediate staging can reserve fast-tier capacity long before retrieval begins. We present TempoKV, a timing-aware resource-commitment layer that separates early knowledge of reuse from the acquisition of staging resources. It records reusable-KV hits as metadata-only claims and requests commitment when the runtime-estimated time until retrieval falls to the storage-estimated time needed to make KV resident and protected against eviction. These estimates adapt to runtime progress and staging state, while commitment remains subject to available protected capacity. We implement TempoKV in vLLM and LMCache on an SSD-backed CXL memory device without changing request scheduling. Across two models and three prefix cache ratios, TempoKV reduces protected fast-tier byte-time per request by 63-91% versus immediate staging while retaining much of the serving benefit of advance staging. In a fast-tier capacity sweep, output throughput and p95 time to first token (TTFT) remain nearly unchanged as capacity decreases from 100 to 25 GiB. Compared with unmodified LMCache's Device-DAX L1 configuration, TempoKV reduces p95 TTFT by up to 48.0% and increases output throughput by up to 27.8%.