执行是否需要目标KV保真度?一种用于LLM服务的混合保真度KV运行时
Does Execution Require Target KV Fidelity? A Mixed-Fidelity KV Runtime for LLM Serving
浏览论文内容
中文总结 AI 辅助
针对LLM服务中KV缓存内存压力问题,提出混合保真度KV运行时ElasticKV,通过紧凑中间状态和压力感知管理,在高并发下显著降低TTFT并保持生成质量。
中文摘要 AI 辅助
大语言模型(LLM)服务日益受到键值(KV)缓存所消耗的GPU内存的制约。现有的压缩、驱逐和卸载技术缓解了这一压力,但服务运行时通常仅将配置的目标KV表示视为可执行状态。在内存压力下,这种仅目标契约可能将KV短缺转化为请求停滞和抢占。我们提出了ElasticKV,一种基于观察(即目标保真度无需门控执行)构建的混合保真度KV运行时。ElasticKV引入了一种紧凑的中间KV状态,使保真度成为运行时管理的执行属性。为了在分页服务运行时中实现这种状态,ElasticKV结合了(i)一种对结构布局,将保真度降低转化为可复用的GPU容量,(ii)一种双模式注意力后端,直接消费紧凑状态同时保留原生仅目标路径,以及(iii)压力感知的保真度管理,根据内存压力调整KV保真度。我们在多样化工作负载、模型家族和规模以及GPU平台上的广泛评估证明了ElasticKV的有效性和通用性。在高并发下,与vLLM相比,ElasticKV实现了3.8-4.0倍更低的首次令牌时间(TTFT)和9.1倍更低的P90 TTFT,同时保持了生成质量。
英文摘要
Large language model (LLM) serving is increasingly constrained by the GPU memory consumed by key-value (KV) caches. Existing compression, eviction, and offloading techniques alleviate this pressure, but serving runtimes typically treat only the configured target KV representation as execution-ready. Under memory pressure, this target-only contract can turn KV shortage into request stalls and preemptions. We present ElasticKV, a mixed-fidelity KV runtime built on the observation that target fidelity need not gate execution. ElasticKV introduces a compact intermediate KV state, making fidelity a runtime-managed execution property. To realize this state in a paged serving runtime, ElasticKV combines (i) a pair-structured layout that turns fidelity reduction into reusable GPU capacity, (ii) a dual-mode attention backend that directly consumes the compact state while preserving the native target-only path, and (iii) pressure-aware fidelity management that adapts KV fidelity to memory pressure. Our extensive evaluation across diverse workloads, model families and scales, and GPU platforms demonstrates the effectiveness and generality of ElasticKV. Under high concurrency, ElasticKV achieves 3.8-4.0$\times$ lower time-to-first-token (TTFT) and 9.1$\times$ lower P90 TTFT than vLLM while preserving generation quality.
发表机构
- University of Melbourne(墨尔本大学)
机构由 AI 辅助整理,请以论文原文为准。