面向智能体AI服务中边缘LLM推理的端到端时延最小化与负载均衡请求调度
End-to-End Latency-Minimizing and Load-Balanced Request Scheduling for Edge LLM Inference in Agentic AI Services
- Concordia University(康考迪亚大学)
- Hang Seng University of Hong Kong(香港恒生大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对智能体AI服务中边缘LLM推理的请求调度问题,提出LYREO框架,通过跨时隙推理模型和Lyapunov优化实现端到端时延最小化与负载均衡,仿真验证其优于基线方案。
AI中文摘要:
基于大语言模型(LLM)的智能体AI服务对低时延推理的需求日益增长,这推动了LLM在分布式边缘服务器上的部署。然而,异构的通信和计算能力,加上动态演变的推理状态,使得每个到达请求的边缘服务器选择随时间变化,且跨时隙紧密耦合。本文研究了一种面向边缘LLM推理的在线请求调度框架,该框架联合最小化长期平均端到端时延,并调节异构边缘服务器之间的工作负载分布。在此背景下出现了两个主要挑战。首先,传统时延模型无法准确捕捉多阶段LLM执行的细粒度动态。其次,调度决策的时延结果仅在请求完成后才能观察到,这使得即时决策评估变得困难。为应对这些挑战,我们开发了一个跨时隙推理模型,该模型捕捉每个多样化请求的传输、预填充、迭代级解码和键值(KV)缓存演变,并通过KV缓存内存-时间消耗指标来表征服务器工作负载。我们提出了LYREO方法,该方法通过Lyapunov优化变换长期负载均衡约束,并采用基于序列的回报预测的奖励再分配,将延迟结果转换为及时的学习信号,以用于更早的决策。在各种配置下的模拟表明,与代表性的基于学习和启发式基线方案相比,LYREO始终实现更低的时延和更均衡的负载分布。
英文摘要:
Large language model (LLM)-powered agentic AI services increasingly demand low-latency inference, motivating the deployment of LLMs across distributed edge servers. However, heterogeneous communication and computing capabilities, together with dynamically evolving inference states, make the edge server selection for each incoming request time-varying and tightly coupled across slots. In this paper, we investigate an online request scheduling framework for edge LLM inference that jointly minimizes long-term average end-to-end latency and regulates workload distribution across heterogeneous edge servers. Two main challenges arise in this context. First, conventional latency models cannot accurately capture the fine-grained dynamics of multi-stage LLM execution. Second, the latency consequence of a scheduling decision is observed only after request completion, making immediate decision evaluation difficult. To address these challenges, we develop a cross-slot inference model that captures transmission, prefill, iteration-level decoding, and key-value (KV) cache evolution for each diverse request, and characterize server workload through a KV cache memory-time consumption metric. We propose the LYREO approach that transforms the long-term load-balancing constraint via Lyapunov optimization and employs reward redistribution with sequencebased return prediction to convert delayed outcomes into timely learning signals for earlier decisions. Simulations under various configurations demonstrate that LYREO consistently achieves lower latency and more balanced load distribution than representative learning-based and heuristic baseline schemes.