GraniKV:面向具有长共享前缀的多智能体系统的非对称粒度KV缓存分页
GraniKV: Asymmetric Granularity KV-Cache Paging for Multi-Agent Systems with Long Shared Prefix
- Seoul National University(首尔大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
GraniKV是首个将非对称分页粒度应用于生产分页服务引擎KV缓存的系统,通过将共享前缀和后缀分别分配至不同存储池,在多智能体服务场景下显著提升了输出令牌吞吐量。
AI中文摘要:
现有的分页服务引擎对KV缓存采用统一的分页粒度,然而多智能体工作负载的两个区域具有相反的存储需求:长共享前缀需要连续性,而每个请求的后缀需要细粒度分配。我们提出了GraniKV,这是一种KV缓存层,它将共享前缀分配到连续的HOT池,将后缀分配到令牌级的COLD池,并结合了一个逐步调度器,该调度器针对每种场景(计算受限、内存受限或通信受限)在双后端中选择合适的后端。据我们所知,GraniKV是第一个将非对称分页粒度应用于生产分页服务引擎KV缓存的系统。在共享令牌数Lp=16K时,GraniKV在Llama-3.1-8B/TP=1、Qwen-2.5-14B/TP=2和Qwen-2.5-32B/TP=4上的输出令牌吞吐量分别达到生产基线的2.16倍、1.98倍和1.57倍。增益可分解:级联注意力集成在饱和状态下贡献了大部分;非对称存储层使端到端性能提升了1.05至1.15倍,且正是该层使得批量GEMM前缀后端成为可能。在具有不同长度不同提示的异构多智能体服务场景中,增益来源发生反转:GraniKV维持1.95倍的性能,而批量全局级联则降至同等水平——仅存储层就承载了该论文所针对场景的性能优势。
英文摘要:
Production paged-serving engines apply uniform paging granularity to the KV cache, even though the two regions of a multi-agent workload have opposite storage requirements: a long shared prefix demands contiguity, while the per-request suffix demands fine-grained allocation. We present \textbf{GraniKV}, a KV-cache layer that allocates the shared prefix in a contiguous HOT pool and the suffix in a token-level COLD pool, combined with a per-step dispatcher which selects the appropriate backend among dual backends for each regime (compute-, memory-, or communication-bound). To the best of our knowledge, GraniKV is the first system to apply asymmetric paging granularity to the KV cache of a production paged-serving engine. At $L_p{=}16$\,K shared tokens GraniKV reaches $\mathbf{2.16\times}$, $\mathbf{1.98\times}$, and $\mathbf{1.57\times}$ output-token throughput over the production baseline on Llama-3.1-8B/TP=1, Qwen-2.5-14B/TP=2, and Qwen-2.5-32B/TP=4. The gain decomposes: cascade attention integration contributes the majority at saturation; the asymmetric storage layer adds $1.05$--$1.15\times$ end-to-end while being what makes the batched-GEMM prefix backend possible at all. Under heterogeneous multi-agent serving with \emph{distinct} prompts of different lengths, the attribution inverts: GraniKV sustains $\mathbf{1.95\times}$ while batch-global cascade collapses to parity --- the storage layer alone carries the win in the regime that motivates the paper.