发表机构
University of Luxembourg; National University of Singapore(卢森堡大学; 新加坡国立大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对6G网络无线资源管理的动态性与约束难题,提出融入OFDM领域知识的两阶段混合学习框架,可扩展MARL算法分解问题,仿真显示其吞吐量提升达40%且约束满足率100%,性能优于现有基线。
AI 中文摘要
无线资源管理(RRM)是6G无线网络中的一项基础挑战,尤其在动态用户需求、小区间干扰和异构QoS约束的场景下更为突出。由于缺乏全系统状态信息、动态条件和信令延迟,集中式优化解决方案在实践中往往不可行,这使得基于分布式学习的方法颇具吸引力。然而,传统多智能体强化学习(MARL)在这类高度动态的环境中面临可扩展性和约束满足方面的困难。我们提出一种可扩展的两阶段混合学习框架来解决RRM挑战,该框架将正交频分复用(OFDM)领域知识明确融入MARL流程。在我们提出的两阶段学习框架中,第一阶段基于信道统计信息分配最小资源以满足用户的QoS要求,避免了在严格QoS约束下训练策略的低效性;随后,第二阶段采用多智能体系统来最优分配剩余资源,以最大化系统吞吐量。通过利用结构化干扰信息,我们提出一种可扩展的MARL算法,该算法将原始学习问题分解为可独立处理的小子问题,从而在不降低性能的情况下避免了动作空间的指数级增长。在具有50MHz带宽和不同 numerology( numerology为5G/6G中无线帧结构参数集的专有术语,保留原名)的现实场景中进行的大量仿真表明,我们的方法显著优于现有的优化和学习基线,提供了高达40%的吞吐量提升,且在资源使用最少的情况下实现了100%的约束满足。
英文摘要
Radio Resource Management (RRM) is a fundamental challenge in 6G wireless networks, particularly under dynamic user demands, inter-cell interference, and heterogeneous QoS constraints. Centralized optimization solutions are often infeasible in practice due to the lack of system-wide state information, dynamic conditions, and signaling delays, making distributed learning-based approaches attractive. However, conventional multi-agent reinforcement learning (MARL) struggles with scalability and constraint satisfaction in such highly dynamic environments. We introduce a scalable two-phase hybrid learning framework to address the RRM challenges where orthogonal-frequency division multiplexing (OFDM) domain knowledge is explicitly incorporated into the MARL pipeline. In our proposed two-phase learning framework, the first phase allocates a minimum resource to satisfy users' QoS requirements based on channel statistics, avoiding the inefficiencies of training policies under hard QoS constraints. Subsequently, a multiagent system is employed in the second phase to optimally allocate the remaining resources for system throughput maximization. By exploiting the structural interference information, we propose a scalable MARL algorithm which decomposes the original learning problem into smaller subproblems that can be handled independently, thereby avoiding exponential growth of the action space without compromising performance. Extensive simulations in realistic scenarios with 50MHz bandwidth and different numerologies show that our method significantly outperforms existing optimization and learning baselines, offering up to 40% throughput improvement, and 100% constraint satisfaction with minimal resource usage.