AI 中文总结
针对解聚式智能体LLM服务,提出SLO感知资源分配框架SARA,基于排队论建模各阶段,在成本约束下最大化有效吞吐量,平均提升26.6%。
AI 中文摘要
大型语言模型(LLM)的最新进展正推动通过云和边缘基础设施为移动用户提供多模态和智能体服务,其中长上下文工作负载对推理延迟提出了严峻挑战。现有的解聚式LLM服务系统主要依赖硬件剖析、配置枚举或启发式调度,在成本效益型资源分配方面提供的分析指导有限。本文提出SARA,一个面向解聚式智能体LLM服务系统的服务等级目标(SLO)感知资源分配框架,该框架在部署成本约束和一系列基于分位数的SLO约束下最大化有效吞吐量。利用排队论,我们首先将预填充、KV缓存传输和解码阶段分别建模为M/G/k队列、M/G/1队列和广义生灭过程。分析表明,预填充和解码阶段分别主要受计算能力和高带宽内存(HBM)资源的限制。基于这些数学模型,我们进一步推导了轻尾和重尾工作负载下不同阶段服务等级指标的可处理尾部行为。这些特征明确地将工作负载、模型架构和硬件参数映射到分阶段SLO约束和最小资源需求。最后,我们开发了一个有效的资源分配框架,以在有限成本预算下最大化系统有效吞吐量。仿真和硬件结果表明,所提出的框架能准确预测分阶段SLO,平均误差低于5%,并且在相同部署成本下,系统有效吞吐量平均比最先进的基线方法提高26.6%。
英文摘要
Recent advances in large language models (LLMs) are driving the emergence of multi-modal and agentic services for mobile users through cloud and edge infrastructures, where long-context workloads pose daunting challenges for inference latency. Existing disaggregated LLM serving systems largely rely on hardware profiling, configuration enumeration, or heuristic scheduling, offering limited analytical guidance for cost-efficient resource allocation. In this paper, we propose SARA, a Service level objectives (SLOs)-Aware Resource Allocation framework for disaggregated agentic LLM serving systems, which maximizes goodput under a deployment cost constraint and a series of quantile-based SLO constraints. By capitalizing on queuing theory, we first model the prefill, KV cache transfer, and decode stages as an M/G/k queue, an M/G/1 queue, and a generalized birth-death process, respectively. The analysis reveals that the prefill and decode stages are dominantly limited by computational capacity and high-bandwidth memory (HBM) resources, respectively. With these mathematical models, we further derive tractable tail behaviors of different stage-wise service level metrics for both light- and heavy-tailed workloads. These characterizations explicitly map workload, model architecture, and hardware parameters to stage-wise SLO constraints and minimum resource requirements. Finally, we develop an effective resource allocation framework to maximize system goodput under limited cost budgets. Simulation and hardware results demonstrate that the proposed framework accurately predicts the stage-wise SLO with mean errors below 5%, and improves system goodput by 26.6% on average over state-of-the-art baseline methods under the same deployment cost.
Comments17 pages, 12 figures