arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

DRLM:边缘环境中基于深度强化学习的大语言模型查询编排框架

DRLM: Deep Reinforcement Learning-Based LLM Query Orchestration in Edge Environments

Reza Farahani, Zoha Azimi Ourimi, Mario Colosi, Lauri Loven, Christian Timmerer, Schahram Dustdar

arXiv 2609.00442首次发表:更新:

发表机构

TU Wien; University of Klagenfurt; University of Messina; University of Oulu(维也纳工业大学; 克拉根福大学; 墨西拿大学; 奥卢大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对边缘环境中LLM查询编排的异构性挑战,提出基于深度强化学习的DRLM框架,通过轻量预测器与因子化PPO智能体实现编排,在64节点集群上较基线方法显著降低延迟且准确率损失小。

AI 中文摘要

大语言模型(LLM)服务日益处理具有不同延迟、准确率和资源需求的异构查询。虽然边缘部署可缩短响应时间,但设备的异构性以及模型家族、参数规模和量化水平的多样性,使得高效的LLM查询编排颇具挑战。本文介绍了边缘环境中基于深度强化学习的LLM查询编排框架DRLM,它集成了两个轻量预测器:(i)将查询映射到语义类别并推断模型性能的类别条件质量估计器;(ii)跨模型-设备配置估计推理时间的特征驱动延迟预测器。这些预测结果结合系统状态,输入到一个因子化近端策略优化(PPO)智能体,该智能体执行感知状态的编排决策。为支持数据驱动的编排,我们构建了一个大规模基准数据集,包含223835个测量值,覆盖1258个查询、6个查询类别、8个模型家族(32个部署实例)、5个量化水平以及异构边缘设备。在64节点边缘集群上的评估,与3个基线方法和2个最先进方法的对比显示,DRLM可将推理延迟最多降低51%,排队延迟最多降低67%,同时准确率损失最多仅8%;在 workload 增加时,其延迟改善可达61.4%,展现出稳健且稳定的编排能力。

英文摘要

Large language model (LLM) services increasingly process heterogeneous queries with diverse latency, accuracy, and resource requirements. While edge deployment reduces response time, the heterogeneity of devices and the diversity of model families, parameter scales, and quantization levels make efficient LLM query orchestration challenging. This paper introduces DRLM, a Deep Reinforcement Learning-based LLM query orchestration framework in edge environments. DRLM integrates two lightweight predictors: (i) a class-conditioned quality estimator that maps queries to semantic categories and infers model performance, and (ii) a feature-driven latency predictor that estimates inference time across model-device configurations. These predictions, combined with system state, feed a factorized Proximal Policy Optimization (PPO) agent that performs state-aware orchestration decisions. To enable data-driven orchestration, we construct a large-scale benchmarking dataset with 223 835 measurements spanning 1258 queries, 6 query classes, 8 model families (32 deployed instances), 5 quantization levels, and heterogeneous edge devices. Evaluation on a 64-node edge cluster and comparison with three baselines and two state-of-the-art methods show that DRLM reduces inference latency by up to 51% and queuing delay by up to 67 %, while incurring at most 8% accuracy loss. It improves latency under increasing workloads up to 61.4%, demonstrating robust and stable orchestration.

Comments6 pages, 7 figures, 1 table, accepted for publication in Globecom 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑