arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.16336cs.ARcs.DCcs.LG

超越二元优先级:面向大语言模型服务的多层级SLA调度

Beyond Binary Priorities: Multi-Tier SLA Scheduling for Large Language Model Serving

Anders Vestrum, Arya Raeesi, Hanna Roed

首次发表
浏览论文内容

中文总结 AI 辅助

本研究扩展Llumnix的优先级模型支持任意层级,在Vidur模拟器中对比多基线,发现四层优先级调度成本效益最优,10层时仍能维持收益且无尾部延迟崩溃。

中文摘要 AI 辅助

现代大语言模型(LLM)服务部署必须同时满足不同用户层级的异构服务水平目标(SLO),涵盖从延迟敏感的API调用到后台批处理的各类场景。Llumnix是一款支持动态迁移的多实例LLM推理调度器,它通过统一的“空闲度”指标实现负载均衡、碎片整理、优先级管理和自动扩缩容。但Llumnix的优先级模型仅支持高、普通两个层级,这种抽象过于粗糙,无法表达生产部署中常见的更丰富的SLA类别。本研究将Llumnix的优先级模型扩展为支持任意数量的层级,并使用高保真LLM推理模拟器Vidur,在均匀分布、高斯分布、企业级分布三种现实优先级分布下评估该扩展的效果。我们在Vidur的分层调度框架内实现了带指数衰减的层级余量、层级感知的调度顺序,以及完整的Llumnix迁移流水线。我们将扩展后的调度器与INFaaS(全局路由基线)、vLLM、Orca、Sarathi-Serve(每副本基线)进行对比,优先级层级范围为1到10。实验表明,四个优先级层级能实现最佳的成本效益权衡,与INFaaS相比,预填充阶段平均加速比最高达8.3倍,端到端P99加速比最高达3.1倍,单位延迟成本降低46%至68%,同时保持各层级间的SLO差异化。我们进一步证明,该系统在10个优先级层级下仍能维持这些收益,且不会出现尾部延迟崩溃,开销集中在预填充阶段。

英文摘要

Modern LLM serving deployments must simultaneously satisfy heterogeneous service-level objectives (SLOs) across a diverse population of user tiers, ranging from latency-critical API calls to background batch processing. Llumnix introduced a dynamic, migration-capable multi-instance scheduler for LLM inference that achieves load balancing, defragmentation, prioritization, and auto-scaling through a unified "freeness" metric. However, Llumnix's priority model is restricted to two levels (high and normal), an abstraction too coarse to express the richer SLA classes common in production deployments. In this work, we extend Llumnix's priority model to support an arbitrary number of tiers and evaluate the effects of this extension under three realistic priority distributions (uniform, Gaussian, enterprise) using Vidur, a high-fidelity LLM inference simulator. We implement per-tier headroom with exponential decay, tier-aware dispatch ordering, and the full Llumnix migration pipeline inside Vidur's hierarchical scheduling framework. We compare our extended scheduler against INFaaS (global routing baseline), vLLM, Orca, and Sarathi-Serve (per-replica baselines), sweeping priority levels from 1 to 10. Our experiments demonstrate that four priority tiers yields the best cost-effectiveness tradeoff, achieving prefill mean speedups of up to 8.3x and end-to-end P99 speedups of up to 3.1x over INFaaS with cost-per-latency improvements of 46 to 68%, while preserving strong SLO differentiation across tiers. We further show that the system sustains these gains at 10 priority levels without tail latency collapse, with overhead concentrated in the prefill phase.

发表机构

  • UC Berkeley(加州大学伯克利分校)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑