arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.34555cs.LG

PulseInfer:面向高效长上下文LLM解码的以I/O为中心的稀疏KV缓存卸载

PulseInfer: I/O-Centric Sparse KV Cache Offloading for Efficient Long-Context LLM Decoding

发表机构华中科技大学 · 优刻得科技股份有限公司
查看机构详情
  • Huazhong University of Science and Technology(华中科技大学)
  • UCloud Technology Co., Ltd.(优刻得科技股份有限公司)

机构由 AI 辅助整理,请以论文原文为准。

Qiuyang Zhang, Kai Zhou, Kai Lu, Haocheng Lu, Jian Zhou, Yuanpeng Su, Kun Bao, Jiguang Wan, Fei Wu

首次发表
浏览论文内容

中文总结 AI 辅助

PulseInfer提出以I/O为中心的稀疏KV缓存卸载系统,通过可中断逐层调度、IO自适应卸载准入和SoloHead稀疏选择等技术,解决长上下文解码的召回I/O瓶颈,将吞吐量提升至4.7倍并降低TPOT达76%。

中文摘要 AI 辅助

长上下文LLM服务日益受限于解码阶段,其中大型KV缓存限制了批处理大小并使GPU利用率不足。稀疏KV缓存卸载通过将大部分历史KV块存储在CPU DRAM中,并仅按需召回选定的块,从而扩大了有效容量。然而,我们发现现有的卸载系统将瓶颈转移到了CPU-GPU召回I/O上:召回量在各层、解码步骤和请求之间差异很大,而按头稀疏选择将召回碎片化为许多小的PCIe传输。本文提出了PulseInfer,一个以I/O为中心的稀疏KV缓存卸载系统。PulseInfer通过可中断的逐层调度隐藏可变的召回延迟,通过IO自适应卸载准入调整卸载决策,并使用SoloHead稀疏选择和聚集-分散I/O引擎合并碎片化传输。在SGLang上实现后,PulseInfer将解码吞吐量相比SGLang提高了最多4.7倍,相比最佳现有卸载基线提高了2.6倍,同时将TPOT降低了最多76%,并保持了近乎无损的准确性。

英文摘要

Long-context LLM serving is increasingly bottlenecked by decode, where large KV caches limit batch size and underutilize GPUs. Sparse KV cache offloading expands effective capacity by storing most historical KV blocks in CPU DRAM and recalling only selected blocks on demand. However, we find that existing offloading systems shift the bottleneck to CPU-GPU recall I/O: recall volume varies widely across layers, decode steps and requests, while headwise sparse selection fragments recalls into many small PCIe transfers. This paper presents PulseInfer, an I/O-centric sparse KV cache offloading system. PulseInfer hides variable recall latency with interruptible layer-wise scheduling, adapts offloading decisions with IO-Adaptive Offloading Admission, and coalesces fragmented transfers using SoloHead sparse selection and a gather-scatter I/O engine. Implemented on SGLang, PulseInfer improves decode throughput by up to 4.7x over SGLang and 2.6x over the best existing offloading baseline, while reducing TPOT by up to 76% and preserving near-lossless accuracy.

↑