arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

用于大语言模型服务的弹性键值缓存:一种工作回收机制,以及分块预填充为何已能缩小差距

Elastic KV Cache for LLM Serving:A Working Reclamation Mechanism, and Why Chunked Prefill Already Closes the Gap

Sathishkumar Sivashanmugam

arXiv 2608.23658首次发表:更新:

发表机构

Amazon Web Services(亚马逊网络服务)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出弹性KV缓存机制,测试发现分块预填充已能缩小预留内存回收的性能差距,明确其适用场景并开源弹性VMM分配器。

AI 中文摘要

大语言模型(LLM)服务引擎在启动时会一次性确定其键值(KV)缓存的大小,永久预留出最坏情况下预填充激活所需的内存空间。在以解码为主的阶段,该预留内存处于闲置状态,但由于它恰好是大型预填充所需的内存,因此无法被分配给KV池。我们研究该预留内存是否可被回收,并构建了一种机制对此进行测试。我们的弹性KV缓存会在解码阶段将预留内存借给KV池,在下一次预填充之前再将其收回,该机制由调度器对下一批任务的前瞻视图驱动。它是CUDA虚拟内存路径上的纯用户空间实现:每一层将两个物理句柄映射到一个连续的虚拟地址范围,因此注意力内核无需修改,也不需要驱动程序补丁。它的取消提交耗时几毫秒,重新提交耗时几十毫秒,可与CUDA图和前缀缓存兼容,且从未触发内存不足事件。若采用相同内存的静态提交则存在风险,在预填充突发时会崩溃,因此动态切换是必要的。在构建完该机制后,我们对其依赖的前提进行了测试,并得到了一个确定的否定结果:仅当较小的预填充块大小会严重损害预填充延迟时,该机制才有收益。在向实际解码负载中注入长提示的受控实验中,该损害很小(8192和32768 token的块大小之间,首次 token 的中位时间差异约为1%),因为预填充受计算限制,而解码每步仅消耗每个序列约1个token。简单降低max_num_batched_tokens比该控制器能回收更多KV内存,且延迟几乎相同。在张量并行(TP)下,该预留内存的占比会被稀释,从TP1时占KV内存的16%降至TP4时的2.7%。我们明确了回收该预留内存仍可能有帮助的场景,并将该机制作为可复用的用户空间弹性VMM分配器开源发布。

英文摘要

An LLM serving engine sizes its key-value (KV) cache once, at startup, permanently setting aside a reserve for the worst-case prefill activation. During decode-dominant phases that reserve sits idle, yet it cannot be handed to the KV pool because it is exactly the memory a large prefill needs. We ask whether this reserve is reclaimable, and build a mechanism to test it. Our elastic KV cache lends the reserve to the KV pool during decode and returns it before prefill, driven by the scheduler's one-step-ahead view of the next batch. It is pure userspace on the CUDA virtual-memory path: two physical handles mapped into one contiguous virtual range per layer, so the attention kernel is unchanged and no driver patch is required. It decommits in a few milliseconds and recommits in tens of milliseconds, works with CUDA graphs and prefix caching, and never triggers an out-of-memory event. A static commit of the same memory is unsafe, crashing on prefill bursts, which makes the dynamic toggle necessary. Having built the mechanism, we test the premise it rests on and report an honest negative result. It only pays off if a small prefill chunk size badly hurts prefill latency. In a controlled experiment injecting long prompts into a live decode load, that penalty is small (median time-to-first-token differs by about 1% between chunk sizes of 8192 and 32768 tokens), because prefill is compute bound and decode consumes only about one token per sequence per step. Simply lowering max_num_batched_tokens recovers more KV than the controller does, at nearly equal latency. The reserve also dilutes under tensor parallelism, from 16% of KV at TP1 to 2.7% at TP4. We state precisely when reclaiming the reserve could still help, and release the mechanism as a reusable userspace elastic-VMM allocator.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑