GroupKV:面向长上下文扩散LLM推理的分层KV缓存管理
GroupKV: Hierarchical KV Cache Management for Long-Context Diffusion LLM Inference
浏览论文内容
中文总结 AI 辅助
针对长上下文扩散LLM推理中KV缓存膨胀与卸载开销问题,提出GroupKV分层缓存管理系统,通过组级稀疏选择、跨层一致性预取及流式预填充,实现最高48倍上下文长度扩展和3.73倍性能提升。
中文摘要 AI 辅助
扩散大语言模型(dLLMs)正成为一种有前景的生成范式,与自回归解码互补。在长上下文场景中,KV缓存膨胀和卸载传输开销已成为推理系统的主要瓶颈。同时,dLLMs中周期性的全序列重计算和局部令牌更新使得KV生命周期更加动态,这使缓存管理和预取调度复杂化,同时使重量级的令牌级索引或聚类方案在解码过程中难以有效摊销。为应对这些挑战,我们提出了\ extsc{GroupKV},一种用于长上下文dLLM推理的轻量级分层KV缓存管理系统。我们观察到,在块式解码下,同一生成块内的令牌倾向于访问高度重叠且空间集中的上下文区域,这使得组级稀疏选择有效。基于这一观察,\ extsc{GroupKV}将上下文划分为连续组,并执行从粗到细的稀疏选择。\ extsc{GroupKV}进一步利用跨层一致性实现预测性预取,并引入陈旧性校正机制以在动态KV更新下保持缓存一致性。此外,\ extsc{GroupKV}采用流式预填充以减少预填充期间的峰值内存消耗。实验表明,在受限GPU内存下,\ extsc{GroupKV}将最大可服务上下文长度扩展了高达$48.00\ imes$,在基于卸载的长上下文设置中将端到端推理性能提升了高达$3.73\ imes$,并保持了有竞争力的任务准确性。
英文摘要
Diffusion large language models (dLLMs) are emerging as a promising generative paradigm that complements autoregressive decoding. In long-context settings, KV cache bloat and offloading transfer overhead have become primary bottlenecks in inference systems. Meanwhile, the periodic full-sequence recomputation and localized token updates in dLLMs make the KV lifecycle substantially more dynamic, complicating cache management and prefetch scheduling while making heavyweight token-level indexing or clustering schemes harder to amortize effectively during decoding. To address these challenges, we present \textsc{GroupKV}, a lightweight hierarchical KV cache management system for long-context dLLM inference. We observe that under block-wise decoding, tokens within the same generation block tend to access highly overlapping and spatially concentrated context regions, making group-level sparse selection effective. Building on this observation, \textsc{GroupKV} partitions the context into contiguous groups and performs coarse-to-fine sparse selection. \textsc{GroupKV} further exploits cross-layer consistency to enable predictive prefetching, and incorporates a staleness correction mechanism to maintain cache coherence under dynamic KV updates. Additionally, \textsc{GroupKV} adopts streaming prefill to reduce peak memory consumption during prefilling. Experiments show that \textsc{GroupKV} extends the maximum serviceable context length by up to $48.00\times$ under constrained GPU memory, improves end-to-end inference performance by up to $3.73\times$ in offload-based long-context settings, and maintains competitive task accuracy.
发表机构
- Institute of Software, Chinese Academy of Sciences(中国科学院软件研究所)
- University of Chinese Academy of Sciences(中国科学院大学)
- Nanjing University of Science and Technology(南京理工大学)
机构由 AI 辅助整理,请以论文原文为准。