arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.29160cs.DC

共享LLM服务的跨模型自动缩放

Cross-Model Autoscaling for Shared LLM Serving

Xin Zhang, Xianyan Xie, Zhen He, Xijin Yin, Xingtong Lin, Bangbo Liang, Zequn Cheng, Peihao Huang, Guo Chen

首次发表
浏览论文内容

中文总结 AI 辅助

提出TRE框架,通过Token服务份额信号和跨模型容量协调,在固定GPU预算下优化共享LLM服务的延迟SLO,显著降低P95/P99延迟。

中文摘要 AI 辅助

多模型LLM服务正朝着共享MaaS集群发展,其中共同托管的模型在固定的GPU预算下竞争,同时每个模型经历时变需求并必须满足自身的延迟SLO。现有的LLM自动缩放器在很大程度上仍是模型本地的:它们的信号暴露本地运行时活动或延迟结果,其扩展或配置决策并不直接决定共享容量应如何在竞争模型之间分配。我们提出了Token服务份额再平衡引擎(TRE),一个用于热切换多模型LLM服务的控制平面框架。TRE引入了Token服务份额(TSS),一种校准的、需求归一化的信号,估计每个活动或排队请求的有效token服务,并在异构模型和SLO类别之间产生可比较的健康分数。在TSS的指导下,TRE在固定GPU预算下协调有界的接收者-捐赠者容量移动:它将快速救援与较慢的再平衡分开,并逐步将活动副本重新分配给校准服务赤字最大的模型。我们在基于Kubernetes的热切换服务栈上实现了TRE,无需修改推理调度器。在七个LLM服务跟踪中,与在同一热切换运行时上运行的最先进的基于KV缓存的自适应自动缩放器相比,TRE将P95端到端延迟降低了11.9-79.0%,P99延迟降低了12.5-72.6%。这些收益在定向压力测试和生产衍生的对话/代码跟踪中均保持,其中TRE分别将P95/P99延迟降低了50.8/63.7%和79.0/72.6%。这些结果表明,有效的热切换自动缩放不仅需要快速的副本激活,还需要校准的服务赤字信号和协调的跨模型容量仲裁。我们的代码和工件可在该https URL获取。

英文摘要

Multi-model LLM serving is moving toward shared MaaS clusters, where co-hosted models compete for a fixed GPU budget while each model experiences time-varying demand and must satisfy its own latency SLO. Existing LLM autoscalers remain largely model-local: their signals expose local runtime activity or delayed latency outcomes, and their scale-up or provisioning decisions do not directly determine how shared capacity should be allocated across competing models. We present the Token-service-share Rebalancing Engine (TRE), a control-plane framework for hot-switched multi-model LLM serving. TRE introduces Token Service Share (TSS), a calibrated, demand-normalized signal that estimates effective token service per active or queued request and yields a comparable health score across heterogeneous models and SLO classes. Guided by TSS, TRE coordinates bounded receiver--donor capacity movement under a fixed GPU budget: it separates fast rescue from slower rebalancing and incrementally reallocates active replicas toward models with the largest calibrated service deficits. We implement TRE on a Kubernetes-based hot-switch serving stack without modifying the inference scheduler. Across seven LLM serving traces, TRE reduces P95 end-to-end latency by 11.9--79.0\% and P99 latency by 12.5--72.6\% compared with a state-of-the-art KV-cache-based reactive autoscaler running on the same hot-switch runtime. The gains hold on both targeted stress probes and production-derived conversation/code traces, where TRE reduces P95/P99 latency by 50.8/63.7\% and 79.0/72.6\%, respectively. These results show that effective hot-switched autoscaling requires not only fast replica actuation, but also calibrated service-deficit signals and coordinated cross-model capacity arbitration. Our code and artifacts are available at https://github.com/zxzx9898/Token-service-share_Rebalancing_Engine.

发表机构

  • Hunan University(湖南大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑