arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

输出令牌长度不确定情况下用于大语言模型服务的鲁棒键值缓存管理

Robust KV Cache Management for LLM Serving under Output Token Length Uncertainty

Jiaming Cheng, Duong The Do, Duong Tung Nguyen

arXiv 2607.16892首次发表:更新:

AI 中文总结

针对大语言模型服务中键值缓存管理难题,提出鲁棒框架,联合优化多项配置,纳入延迟SLO约束。开发DRO公式及算法应对输出令牌长度不确定等问题,分析出关键分位数结构,经评估比基线成本降56%,保持相关性能。

AI 中文摘要

键值缓存内存是部署在GPU集群上的现代大语言模型服务系统的主要瓶颈。一个根本挑战是,在请求到达时必须预留键值缓存,而输出令牌长度在生成完成之前是未知的。预留不足会触发抢占,导致请求终止和重新计算,并产生大量开销;而预留过多则会浪费内存并降低吞吐量。这在内存效率和抢占风险之间形成了核心权衡。我们提出了一个用于大语言模型服务的鲁棒键值缓存管理框架,该框架联合优化GPU并行配置、每个请求类别的键值缓存预留、跨异构服务组的请求路由以及共享提示的前缀缓存。该框架纳入了延迟服务水平协议(SLO)约束,并捕捉了内存分配、吞吐量和排队延迟之间的相互作用。为了解决输出令牌长度的不确定性和工作负载分布的变化,我们开发了一种瓦瑟斯坦分布鲁棒优化(DRO)公式以及一种用于由此产生的混合整数问题的可扩展块坐标下降算法。我们的分析揭示了一种关键的分位数结构,该结构可以自动将预留分位数适应于不同的抢占和内存成本模式,而无需手动调整。对包括BurstGPT、Azure和ShareGPT跟踪在内的生产大语言模型工作负载的跟踪驱动评估表明,与固定分位数预留基线相比,成本降低了56%,同时在不同的运行模式下保持了有竞争力的P99延迟、吞吐量和SLO违规率。

英文摘要

KV cache memory is a primary bottleneck in modern LLM serving systems deployed on GPU clusters. A fundamental challenge is that the KV cache must be reserved upon request arrival, while the output token length remains unknown until generation completes. Under-reservation triggers preemption -- forcing termination and recomputation of requests and incurring significant overhead -- whereas over-reservation wastes memory and reduces throughput. This creates a central trade-off between memory efficiency and preemption risk. We present a robust KV cache management framework for LLM serving that jointly optimizes GPU parallelism configuration, KV cache reservation per request class, request routing across heterogeneous serving groups, and prefix caching for shared prompts. The framework incorporates latency SLO constraints and captures the interaction between memory allocation, throughput, and queueing delay. To address output token length uncertainty and workload distribution shift, we develop a Wasserstein distributionally robust optimization (DRO) formulation together with a scalable block coordinate descent algorithm for the resulting mixed-integer problem. Our analysis reveals a critical fractile structure that automatically adapts reservation quantiles to different preemption and memory cost regimes without manual tuning. Trace-driven evaluation on production LLM workloads, including BurstGPT, Azure, and ShareGPT traces, demonstrates up to 56\% lower cost than fixed-quantile reservation baselines while maintaining competitive P99 latency, goodput, and SLO violation rates across diverse operating regimes.

Comments10 figures, 10 pages

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑