面向高效多租户模型服务的全局仿真引导型动态算子调度
Global Simulation-Guided Dynamic Operator Scheduling for Efficient Multi-Tenant Model Serving
- Xiamen University(厦门大学)
- Shanghai Innovation Institute(上海创新研究院)
- Shanghai Jiao Tong University(上海交通大学)
- East China Normal University(华东师范大学)
- Tsinghua University(清华大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文提出SliceScheduler系统,通过全局映射图、全局仿真器等组件实现算子级调度,在维持SLA违规率低于9%的同时,将多租户LLM服务的令牌吞吐量提升1.10-2.29倍,有效提高GPU利用率。
AI中文摘要:
容器粒度的调度会在容器内留下大量未被利用的短寿命空闲切片;而在服务水平协议(SLA)约束下,重新分配容器过于笨重,难以利用这类细粒度机会,且算子级调度需要实时推理依赖关系、内存安全性以及集群范围的执行动态。本文提出SliceScheduler,一种面向多租户模型服务的动态算子级调度系统,核心思路是暴露集群范围的算子执行状态,并支持对调度决策进行假设推理。SliceScheduler包含四个关键组件:其一,引入全局映射图(GMG),这是一种统一抽象,可捕获算子依赖、张量形状、资源映射和执行状态,提供带有显式资源语义的实时集群范围视图;其二,在GMG之上构建全局仿真器,用于预测候选部署下的算子级执行与内存演变;其三,设计基于增量仿真的调度模块,选择部署方式以利用碎片化空闲切片,同时避免内存违规并维持SLA;其四,开发算子执行器,在GPU上实现调度决策,并协调计算与跨加速器传输。我们将SliceScheduler实现为PyTorch后端,并通过生产轨迹重放进行评估。实验结果显示,与现有方法相比,SliceScheduler将令牌吞吐量提升了1.10至2.29倍,同时将SLA违规率控制在9%以内,证明算子级调度是提升多租户大语言模型(LLM)服务GPU利用率的实用且有效方法。
英文摘要:
Container-granularity scheduling leaves abundant short-lived idle slices within containers unexploited. Reallocating containers is too heavyweight to utilize such fine-grained opportunities under SLA constraints, and operator-level scheduling requires reasoning about dependencies, memory safety, and cluster-wide execution dynamics in real time. In this paper, we present SliceScheduler, a dynamic operator-level scheduling system for multi-tenant model serving. The key idea is to expose cluster-wide operator execution state and enable what-if reasoning over scheduling decisions. SliceScheduler consists of four key components. First, we introduce the Global Mapping Graph (GMG), a unified abstraction that captures operator dependencies, tensor shapes, resource mappings, and execution states, providing a real-time, cluster-wide view with explicit resource semantics. Second, we build a global simulator on top of GMG to predict operator-level execution and memory evolution under candidate placements. Third, we design an incremental, simulation-based scheduling module that selects placements to exploit fragmented idle slices while avoiding memory violations and preserving SLA. Finally, we develop an operator executor that materializes scheduling decisions on GPUs and coordinates computation and cross-accelerator transfers. We implement SliceScheduler as a PyTorch backend and evaluate it using production trace replay. Experimental results show that SliceScheduler improves token throughput by 1.10--2.29$\times$ compared to existing approaches, while maintaining SLA violations within 9\%. SliceScheduler demonstrates that operator-level scheduling is a practical and effective approach to improving GPU utilization for multi-tenant LLM serving.