arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

OasisKV:通过前瞻稀疏预取将解码中KV缓存扩展至HBM之外

OasisKV: Scaling In-Decode KV Cache Beyond HBM with Lookahead Sparse Prefetching

Can Xiao, Sukmin Cho, Junbong We, Zhixiong Niu, Jianyi Cheng, Yiren Zhao, Youngjin Kwon, Yongqiang Xiong, Rui Ma, Junyi Liu

arXiv 2608.08097首次发表:更新:

AI 中文总结

针对LLM推理中HBM容量受限问题,提出以内存为中心的OasisKV系统,通过前瞻稀疏预取技术解耦KV缓存存储与HBM,在精度损失极小的情况下显著提升推理吞吐量并降低内存占用。

AI 中文摘要

大型语言模型(LLM)推理服务日益受到内存而非计算的限制。随着长上下文和长形式推理工作负载愈发普遍,键值(KV)缓存会在LLM的token生成(即解码)阶段主导内存占用和内存流量。尤其,高带宽内存(HBM)容量已成为稀缺且昂贵的资源,严重限制了推理批处理规模和系统吞吐量。本文提出OasisKV,一种以内存为中心的LLM推理系统设计,通过在LLM解码期间将完整KV缓存存储与HBM解耦,缓解HBM容量压力。由于解码时注意力天然具有稀疏性,OasisKV仅在HBM中保留注意力计算所需的最相关token的KV条目。我们观察到,利用投机解码(SD)生成的前瞻token可准确预测未来重要token。OasisKV采用高效的注意力背景流水线识别重要KV块,随后将其从更高容量的内存层(如主机或远程内存)预取,并在下一个解码步骤使用前暂存至HBM。我们基于vLLM实现OasisKV,在2048-token的KV预算下,前瞻预测的准确性足以使精度与完整注意力的差距控制在0.7个点以内。这使OasisKV能将稀疏性转化为吞吐量增益:在推理工作负载上,精度损失0.1个点时,相较于密集vLLM实现1.69倍的吞吐量;在多GPU长上下文服务上,吞吐量最高达2.1倍。在预填-解码解耦场景下,OasisKV实现约2倍于密集方案的吞吐量,同时每个请求的KV量减少6.5-9.7倍,且相较于完整KV传输,解码节点主机内存占用减少2.2-2.6倍。

英文摘要

Large language model (LLM) inference serving is increasingly constrained by memory rather than compute. As long-context and long-form reasoning workloads become more prevalent, the key-value (KV) cache dominates both memory footprint and memory traffic during LLM token generation, i.e., decode. In particular, HBM capacity has become a scarce and costly resource that heavily limits inference batch size and system throughput. This paper presents OasisKV, a memory-centric LLM inference system design that alleviates HBM capacity pressure by decoupling full KV-cache storage from HBM during LLM decoding. Because decode-time attention is naturally sparse, OasisKV keeps only the KV entries of the most relevant tokens in HBMs for attention computation. We observe that future important tokens can be predicted accurately in advance using lookahead tokens drafted by speculative decoding (SD). OasisKV employs an efficient attention background pipeline to identify important KV blocks. They are then prefetched from higher-capacity memory tiers (e.g., host or remote memory) and staged in HBMs before being used in the next decode step. We implement OasisKV based on vLLM. The lookahead prediction is accurate enough to keep accuracy within 0.7 points of full attention under a 2,048-token KV budget. This lets OasisKV turn sparsity into throughput gain: $1.69\times$ over dense vLLM on the reasoning workload at 0.1 points of accuracy loss, and up to $2.1\times$ on multi-GPU long-context serving. Under prefill--decode disaggregation, OasisKV reaches about $2\times$ dense throughput while admitting each request with $6.5$--$9.7\times$ less KV and holding $2.2$-$2.6$ less decode-node host memory than full KV transfer.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑