arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

LMEdge:边缘集群上的QoS感知大语言模型推理编排

LMEdge: QoS-Aware LLM Inference Orchestration on Edge Clusters

Reza Farahani, Zoha Azimi, Mario Colosi, Schahram Dustdar

arXiv 2607.17175首次发表:更新:

AI 中文总结

研究在边缘集群上进行QoS感知的LLM推理编排问题,提出LMEdge服务。通过BILP优化并结合五个轻量级ML模型预测相关指标,设计轻量级启发式方法。在边缘测试平台评估,结果显示其降低延迟、保持准确性、提高资源利用率及服务比率。

AI 中文摘要

大语言模型(LLM)服务越来越多地在边缘基础设施上运行,以实现低延迟和隐私保护的人工智能服务。然而,要在异构且资源受限的边缘设备上高效地服务LLM请求,需要编排机制来共同确定模型配置(系列、大小和量化级别)和执行位置,同时满足用户和系统级的服务质量(QoS)要求。本文介绍了LMEdge,一种跨异构边缘设备动态做出这些决策的QoS感知编排服务。我们将该问题表述为一个二进制整数线性规划(BILP)优化问题,在准确性、网络和资源约束下最小化响应时间。为了实现可扩展的在线调度,我们使用五个轻量级机器学习(ML)模型来预测每个模型大小量化设备组合的特定查询延迟、准确性、资源使用和响应大小,并设计了一种轻量级启发式方法来近似BILP解决方案。我们收集了一个超过59000行的综合基准数据集来训练模型并支持可重复性。在一个基于Kubernetes的具有57个实例和不同查询类别的边缘测试平台上进行的评估表明,与两个基线相比,LMEdge降低了延迟,保持了准确性,提高了资源利用率,并提高了服务比率。

英文摘要

Large language model (LLM) services increasingly operate on edge infrastructure, enabling low-latency and privacy-preserving AI services. However, efficiently serving LLM requests across heterogeneous and resource-constrained edge devices require orchestration mechanisms that jointly determine model configuration (family, size, and quantization level) and execution placement while satisfying user- and system-level quality of service (QoS) requirements. This paper introduces LMEdge, a QoS-aware orchestration service that dynamically makes these decisions across heterogeneous edge devices. We formulate the problem as a binary integer linear programming (BILP) optimization that minimizes response time under accuracy, network, and resource constraints. To enable scalable online scheduling, we employ five lightweight machine learning (ML) models to predict query-specific latency, accuracy, resource usage, and response size for each model-size-quantization-device combination, and design a lightweight heuristic that approximates the BILP solution. We collect a comprehensive benchmarking dataset of over 59000 rows to train models and support reproducibility. Evaluation on a Kubernetes-based edge testbed with 57 instances and diverse query categories shows that LMEdge reduces latency, preserves accuracy, improves resource utilization, and increases serving ratio compared to two baselines.

Comments12 pages, 5 figures, 2 tables, EUROPAR conference paper

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑