arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

耦合有限时域约束下用于无线资源管理的异构多智能体强化学习

Heterogeneous Multi-Agent Reinforcement Learning for Radio Resource Management under Coupled Finite-Horizon Constraints

Yeonseo Jeong, Wonhyeok Ko, Sungweon Hong, Songnam Hong

arXiv 2608.01745首次发表:更新:

发表机构

Hanyang University(汉阳大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对耦合有限时域约束下的无线资源管理难题,提出嵌入李雅普诺夫的异构多智能体强化学习框架HeLyMARL,通过虚拟队列实现预算约束的逐部分时域绑定,仿真中其性能优于多种基准方法。

AI 中文摘要

在密集无线网络中,要实现比例公平性下的吞吐量最大化,需要在硬有限时域能量预算和切换预算的约束下,协同管理用户关联、调度、基站(BS)激活以及切换控制,这就导致了基站侧能量管理与用户侧切换调控之间存在根本的权衡。多智能体强化学习(MARL)是这类分布式序贯控制的自然框架,但将其应用于此场景面临两个难点:一是有限时域预算约束无法在每个时隙进行评估;二是非线性的比例公平效用无法进行原则性的逐时隙分解。我们提出了HeLyMARL,这是一种嵌入李雅普诺夫(Lyapunov)的异构MARL框架,它通过带虚拟队列的漂移加惩罚分解解决了上述两个问题。能量和切换约束压力被直接内化到统一的逐时隙奖励中,从而将受约束的有限时域问题转化为无约束的MARL问题。与两种基于拉格朗日的替代方案的对比显示出时间尺度差异:拉格朗日松弛仅在训练周期内调控约束,而HeLyMARL的虚拟队列则在一个周期内的每个部分时域都对累积预算消耗进行了约束,这是贪心李雅普诺夫控制无法实现的 pacing 保证。仿真结果表明,HeLyMARL是唯一能在整个时域内维持吞吐量-公平性平衡并保证服务不中断的方法,它在预算未提前耗尽的情况下,性能优于传统MARL、基于李雅普诺夫的方法以及受约束MARL基准。

英文摘要

Maximizing throughput under proportional fairness in dense wireless networks requires jointly managing user association, scheduling, base station (BS) activation, and handover control under hard finite-horizon energy and handover budgets, which induces a fundamental tension between BS-side energy management and user-side handover regulation. While multi-agent reinforcement learning (MARL) is a natural framework for such distributed sequential control, its application here faces two difficulties: finite-horizon budget constraints cannot be evaluated at each time slot, and the nonlinear proportional fairness utility admits no principled per-slot decomposition. We propose HeLyMARL, a Lyapunov-embedded heterogeneous MARL framework that resolves both via drift-plus-penalty decomposition with virtual queues. The energy and handover constraint pressures are internalized directly into a unified per-slot reward, converting the constrained finite-horizon problem into an unconstrained MARL problem. Comparison against two Lagrangian-based alternatives reveals a timescale separation: Lagrangian relaxation regulates constraints only across training episodes, whereas the virtual queues of HeLyMARL bound cumulative budget consumption at every partial horizon within an episode, a pacing guarantee beyond the reach of greedy Lyapunov-based control. Simulations show that HeLyMARL is the only method that sustains the throughput-fairness balance together with uninterrupted service throughout the horizon, outperforming conventional MARL, Lyapunov-based, and constrained MARL benchmarks without premature budget exhaustion.

Comments13 pages, 6 figures, 4 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑