arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.20957cs.PF

面向SLO驱动的自动扩缩容的LLM推理服务近似排队模型

An Approximate Queueing Model of LLM Inference Serving for SLO-Driven Autoscaling

Vishakha Ramani, Asser N. Tantawi

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出一种近似排队模型,用于预测LLM推理的TTFT和ITL,并基于该模型实现SLO驱动的自动扩缩容控制器,在H100 GPU集群上有效跟踪负载变化。

中文摘要 AI 辅助

LLM服务器的性能模型支持延迟评估以及针对服务级别目标(SLO)的自动扩缩容控制器设计和推理优化。我们在马尔可夫假设下,用一个易处理的近似排队模型对预填充和解码操作的多路复用执行进行建模。三个参数表征一个模型-加速器对,即每迭代基线开销、每令牌计算成本和每令牌键值(KV)缓存访问成本。该模型将每迭代工作的均值分析与批处理占用的状态相关马尔可夫链相结合,以预测平均首令牌时间(TTFT)和令牌间延迟(ITL)。我们针对测量结果验证了这些预测,并表明这三个参数可以从观测到的延迟中估计出来。在输入和输出长度以及从轻到中等负载的到达率的网格上,Llama-3.1-8B和Qwen2.5-14B在H100 GPU上运行的平均ITL相对误差分别约为5%和8%,相应的TTFT误差为14%和16%。然后,我们实现了一个自动扩缩容控制器,该控制器使用该模型在负载变化时调整推理服务器副本数量。在一个由H100 GPU组成的OpenShift集群上,它在两种延迟目标下跟踪四倍负载爬升,在127个控制周期中有1个未达标,其循环内预测的TTFT中位误差至多为5%,ITL中位误差至多为9%。来自现有自动扩缩容器的解码吞吐量分析器(不考虑延迟目标)在同一控制器和负载下,在128个周期中有27个未达标,同时提供的副本数减少4%和28%。

英文摘要

Performance models of LLM servers support both latency evaluation and the design of controllers for autoscaling against service level objectives (SLOs) and for inference optimization. We model the multiplexed execution of prefill and decode operations with a tractable, approximate queueing model under Markovian assumptions. Three parameters characterize a model-accelerator pair, namely a baseline per-iteration overhead, a per-token compute cost, and a per-token key-value (KV) cache access cost. The model combines a mean-value analysis of per-iteration work with a state-dependent Markov chain for batch occupancy to predict mean time to first token (TTFT) and inter-token latency (ITL). We validate these predictions against measurements and show that the three parameters can be estimated from observed latencies. Over a grid of input and output lengths and arrival rates spanning light to moderate load, the relative error of the average ITL is about 5% for Llama-3.1-8B and 8% for Qwen2.5-14B running on an H100 GPU, and the corresponding TTFT errors are 14% and 16%. We then implement an autoscaling controller that uses the model to adjust inference-server replica counts as the workload changes. On an OpenShift cluster of H100 GPUs it tracks a fourfold load ramp under both latency targets, missing one in 7 of 127 control cycles, and its in-loop predictions carry median errors of at most 5% for TTFT and 9% for ITL. A decode-throughput analyzer from an existing autoscaler, which takes no latency target, misses 27 of 128 cycles under the same controller and load while provisioning 4% and 28% fewer replicas.

发表机构

  • IBM T. J. Watson Research Center(IBM T.J.沃森研究中心)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑