AI 中文总结
该研究针对长上下文LLM推理的内存墙瓶颈,提出轻量级服务系统DualDecoder,通过双token解码流水线预测预取关键KV条目,消除辅助状态开销,使解码吞吐量最高提升2.62倍且保留延迟与模型质量。
AI 中文摘要
长上下文推理正成为现代大语言模型(LLM)服务的基础能力,尤其受新兴智能体(agentic)应用的驱动。但它面临严重的内存墙问题:键值(KV)缓存随上下文长度和请求并发度的增加成比例扩展。现有稀疏KV缓存方法将大部分KV条目卸载到主机内存,仅检索每个解码步骤所需的关键KV条目。然而,这些方法通常会在GPU内存中引入大量用于KV检索管理的辅助状态。我们的测量显示,这些常被忽视的辅助状态会带来显著的内存开销,并在高并发工作负载下成为新的瓶颈。本文提出DualDecoder,一种用于长上下文LLM推理的轻量级服务系统,可实现从主机内存高效检索稀疏KV缓存。我们的核心见解是,解码下一个token所需的关键KV条目可通过前一个推测的token准确预测。这种可预测性使KV检索能够被主动预取并与解码计算重叠,有效消除了辅助状态的GPU内存开销。为高效实现该预取,DualDecoder利用一种新颖的双token解码流水线,以可忽略的计算开销准确识别关键KV条目,并设计了层感知的传输调度,以将KV预取与模型计算重叠,同时设计了层范围的内存管理器以减少GPU运行时缓冲区。实验结果表明,DualDecoder在保持解码延迟和模型质量的同时,比最先进的系统将解码吞吐量提高了2.62倍。
英文摘要
Long-context inference is becoming a fundamental capability for modern LLM serving, especially driven by emerging agentic applications. Yet it faces a severe memory wall that the KV cache scales proportionally with increasing context length and request concurrency. Existing sparse KV cache methods offload most KV entries to host memory and retrieve only the critical KV entries needed by each decoding step. However, they commonly introduce substantial auxiliary states in GPU memory for KV retrieval management. Our measurements show that these often-overlooked auxiliary states introduce significant memory overhead and become a new bottleneck under high-concurrency workloads. In this paper, we present DualDecoder, a lightweight serving system for long-context LLM inference that enables efficient sparse KV cache retrieval from host memory. Our key insight is that the critical KV entries required for decoding the next token can be accurately predicted from the preceding speculated token. This predictability enables KV retrieval to be proactively prefetched and overlapped with decoding computation, effectively eliminating the GPU memory overhead of auxiliary states. To achieve this prefetching efficiently, DualDecoder leverages a novel dual-token decoding pipeline that accurately identifies critical KV entries with negligible computational overhead, and designs a layer-aware transfer schedule to overlap KV prefetching with model computation and a layer-scoped memory manager to reduce the GPU runtime buffer. Experimental results show that DualDecoder improves decoding throughput by up to 2.62$\times$ over state-of-the-art systems while preserving decoding latency and model quality.