PackServe:面向大规模智能体LLM服务的SLO感知请求调度
PackServe: SLO-Aware Request Scheduling for Agentic LLM Serving at Scale
浏览论文内容
中文总结 AI 辅助
PackServe通过白盒延迟预测模型指导请求打包,在保留KVC重用和SLO约束下减少GPU资源消耗,在64块H20 GPU上节省高达24.6%的GPU小时,并在生产集群中降低34.7%资源占用。
中文摘要 AI 辅助
请求调度是大规模集群中服务智能体大语言模型(LLM)工作负载的关键挑战。有效的调度器必须保留跨长共享前缀的键值缓存(KVC)重用,满足令牌级延迟服务级别目标(SLO),并最小化GPU资源占用。现有调度器难以兼顾这些要求:请求合并可能牺牲缓存局部性并增加预填充/解码干扰,从而损害SLO达成率和资源效率。我们提出PackServe,一种旨在降低资源成本同时满足智能体LLM服务延迟SLO的调度器。PackServe使用紧凑白盒模型预测预填充/解码干扰下的延迟。在这些预测的指导下,它将请求打包到更少的服务实例上,同时保留KVC重用和SLO约束,用可用的延迟余量换取更高的每实例吞吐量。在64块NVIDIA H20 GPU上的评估显示,在30毫秒和50毫秒TPOT目标下,PackServe分别比最先进的调度器少使用高达16.8%和24.6%的GPU小时数,同时满足目标TPOT目标。PackServe也已部署在我们包含超过1000块GPU的生产集群中,与原始生产调度器相比,资源占用减少了34.7%。
英文摘要
Request scheduling is a key challenge in large-scale clusters serving agentic large language model (LLM) workloads. An effective scheduler must preserve key-value cache (KVC) reuse across long, shared prefixes, meet token-level latency service-level objectives (SLOs), and minimize GPU resource footprint. Existing schedulers struggle to reconcile these requirements: request consolidation can sacrifice cache locality and increase prefill/decode interference, compromising both SLO attainment and resource efficiency. We present PackServe, a scheduler designed to reduce resource costs while meeting latency SLOs for agentic LLM serving. PackServe uses compact white-box models to predict latency under prefill/decode interference. Guided by these predictions, it packs requests onto fewer serving instances while preserving KVC reuse and SLO constraints, trading available latency headroom for improved per-instance throughput. Evaluation on 64 NVIDIA H20 GPUs shows that PackServe uses up to 16.8% and 24.6% fewer GPU-hours than state-of-the-art schedulers under 30-ms and 50-ms TPOT targets, respectively, while meeting the target TPOT objectives. PackServe has also been deployed in our production cluster comprising over 1000 GPUs, where it reduces the resource footprint by 34.7% compared with the original production scheduler.
发表机构
- The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳))
- Ant Group(蚂蚁集团)
机构由 AI 辅助整理,请以论文原文为准。