OpScale:面向大语言模型服务的算子级资源配置与自动扩缩容
OpScale: Operator-level Provisioning and Autoscaling for LLM Serving
浏览论文内容
中文总结 AI 辅助
针对LLM服务中粗粒度扩缩容的弊端,提出算子级编排框架OpScale,可在满足SLO的同时降低GPU用量、功耗,或在固定成本下提升吞吐量。
中文摘要 AI 辅助
在为大语言模型(LLM)提供服务的云GPU集群中,在满足面向用户的严格服务水平目标(SLO,如首令牌响应时间)的同时实现成本效率,仍是一项核心挑战。自动扩缩容是集群资源管理的关键机制,但针对LLM服务,一个基础的系统设计问题尚未解决:扩缩容的单位应是什么?现有方法主要将整个模型视为单一的扩缩容单位,这种方式简单但无法捕捉推理工作负载的细粒度动态,因此这类粗粒度扩缩容常导致突发需求下的SLO违规,或造成GPU的严重闲置。我们的分析显示,算子存在显著异质性,这表明算子级弹性是可行的扩缩容原语。我们提出OpScale,一个实用的算子级编排框架,涵盖性能分析、资源配置、放置和运行时服务,该框架旨在解决因采用更细粒度操作而产生的高复杂性和空间爆炸问题。基于最多40个A100 GPU和24个GB200 GPU的生产轨迹进行评估,OpScale在实现SLO的同时,可减少多达36.3%的GPU使用量和28%的功耗,或在固定成本预算下实现44%的更高吞吐量。
英文摘要
Achieving cost efficiency while meeting strict user-facing SLOs (e.g., time-to-first-token) remains a fundamental challenge for cloud GPU clusters serving large language models (LLMs). Autoscaling is the key mechanism for cluster resource management, yet a basic system design question is open for serving LLMs: what should be the unit of scaling? Existing approaches primarily treat the entire model as a monolithic scaling unit--simple but unable to capture the fine-grained dynamics of inference workloads. As a result, such coarse-grained scaling often leads to either SLO violations under bursty demand or significant GPU under-utilization. Our characterization reveals substantial operator heterogeneity, exposing operator-level elasticity as a viable scaling primitive. We present OpScale, a practical operator-level orchestration framework of profiling, provisioning, placement, and runtime serving. OpScale is designed to tackle the high complexity and the space explosion problem, arising from operating at this finer granularity. Evaluated with production traces on up to 40 A100s and 24 GB200s, OpScale attains SLOs with up to 36.3% fewer GPUs and 28% less power, or achieves 44% higher throughput under fixed cost budgets.