arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

面向时间敏感型应用的在线流量调度的多智能体强化学习

Multi-Agent Reinforcement Learning for Online Traffic Scheduling in Time-Sensitive Application

Marcos Carvalho, Fatih Temiz, Shavbo Salehi, Melike Erol-Kantarci, Daniel F. Macedo

arXiv 2608.05346首次发表:更新:

发表机构

Universidade Federal de Minas Gerais; School of Electrical Engineering and Computer Science, University of Ottawa(米纳斯吉拉斯联邦大学; 渥太华大学电气工程与计算机科学学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对动态XR驱动的MEC场景中TSN调度的挑战,本文提出基于HAPPO算法的MARL框架,将各TSN队列建模为自主智能体,可降低平均帧等待时间最多26.8%、最坏情况延迟约16.8%。

AI 中文摘要

时间敏感网络(TSN)正日益与移动边缘计算(MEC)集成,以支持扩展现实(XR)等具有严格延迟要求的应用。然而,现有的TSN调度解决方案主要依赖静态优化技术或基于固定流量模式的集中式学习模型,限制了其在动态环境中的有效性。在实践中,MEC环境通常承载多个共存的XR流量流,其特征随时间演变,形成当前调度器无法捕获的复杂队列间依赖关系。应对这些挑战需要自适应的分布式调度机制,能够在不同流量条件下协调多个TSN队列。为此,本文提出一种用于TSN调度的多智能体强化学习(MARL)框架,其中每个TSN队列被建模为自主智能体。采用异构智能体近端策略优化(HAPPO)算法显式建模智能体间的依赖关系,并联合优化各队列的服务交付。仿真结果表明,所提方法可将平均帧等待时间最多降低26.8%,最坏情况延迟降低约16.8%,凸显其在动态XR驱动的MEC场景中的有效性。

英文摘要

Time-sensitive networking (TSN) is increasingly integrated into mobile edge computing (MEC) to support applications with stringent latency requirements, such as extended reality (XR). However, existing TSN scheduling solutions predominantly rely on static optimization techniques or centralized learning models that are based on fixed traffic patterns, limiting their effectiveness in dynamic environments. In practice, MEC environments often host multiple co-located XR traffic flows whose characteristics evolve over time, creating complex inter-queue dependencies that current schedulers fail to capture. Addressing these challenges requires adaptive, decentralized scheduling mechanisms capable of coordinating multiple TSN queues under varying traffic conditions. To this end, this paper proposes a multi-agent reinforcement learning (MARL) framework for TSN scheduling, where each TSN queue is modeled as an autonomous agent. The Heterogeneous-Agent Proximal Policy Optimization (HAPPO) algorithm is employed to explicitly model inter-agent dependencies and jointly optimize service delivery across queues. The simulation results demonstrate that the proposed approach reduces average frame waiting times by up to 26.8% and worst-case delays by approximately 16.8%, highlighting its effectiveness in dynamic XR-driven MEC scenarios.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑