发表机构
The University of British Columbia; Microsoft Azure Research; NVIDIA(不列颠哥伦比亚大学; 微软Azure研究院; 英伟达)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
Cascade是一款LLM服务系统,利用单请求延迟预算协同调度与KV缓存管理,在保留异构请求公平性的同时,使LLM推理服务吞吐量最高提升2.4倍,SLO违规率降低40%。
AI 中文摘要
大型语言模型(LLM)的推理与智能体能力拓展了其应用范围,涵盖从简短交互对话到长耗时计算请求等多种场景。当前LLM服务平台定义了响应延迟的服务水平目标(SLO),但同一服务内的请求在输入长度、生成长度、执行成本及可复用KV缓存状态的可用性方面存在数量级差异。因此,受同一SLO约束的请求具有不同的紧急程度:在扣除执行所需时间后,部分请求有充足的延迟余量,而部分请求几乎无余量。我们将该余量定义为单请求延迟预算,即请求的SLO与预测剩余服务时间的差值。本文提出Cascade,一款LLM服务系统,可基于请求特征、KV缓存状态及当前系统负载估算并持续更新该预算。与仅用截止时间管控请求排序的现有SLO感知调度器不同,Cascade利用单请求预算协同协调内存层级间的请求调度与KV缓存管理:其调度器优先处理剩余预算少的请求,内存管理器则用同一预算决定是否从更深层级恢复或预取非驻留KV状态、将其保留在高带宽内存(HBM)中,或重新计算。通过将排队与数据移动开销导向可吸收该开销的请求,Cascade在保留异构请求类公平性的同时提升了满足SLO的吞吐量。在覆盖三款大型语言模型的生产 traces 上,相较于默认的vLLM先来先服务调度器,Cascade的吞吐量提升最高达2.4倍,SLO违规率降低40%。
英文摘要
The reasoning and agentic capabilities of large language models have expanded the range of applications they support, from short interactive exchanges to long, compute-heavy requests. LLM serving platforms today define response-latency service-level objectives, even though requests within the same service can differ by orders of magnitude in input length, generation length, execution cost, and the availability of reusable KV-cache state. As a result, requests governed by the same service level objective have different urgency: after accounting for the time required to execute them, some have substantial latency headroom while others have almost none. We define this headroom---the difference between a request's service level objective and its predicted remaining service time---as its per-request latency budget. We present Cascade, an LLM serving system that estimates and continuously updates this budget from request characteristics, KV-cache state, and current system load. Unlike prior SLO-aware schedulers that use deadlines to govern request ordering alone, Cascade uses a single per-request budget to jointly coordinate request scheduling and KV-cache management across the memory hierarchy. Its scheduler prioritizes requests with little remaining budget, while its memory manager uses the same budget to decide whether non-resident KV state should be restored or prefetched from a deeper tier, retained in HBM, or recomputed. By directing queueing and data-movement overhead toward requests that can absorb it, Cascade improves SLO-satisfied goodput while preserving fairness across heterogeneous request classes. On production traces across three large language models, Cascade improves goodput by up to2.4x and reduces SLO violations by 40% relative to the default vLLM first-come, first-served scheduler.