arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

可移动天线辅助无蜂窝大规模MIMO系统的多智能体强化学习

Multi-Agent Reinforcement Learning for Movable Antenna-aided Cell-Free Massive MIMO Systems

Bokai Xu, Jiayi Zhang, Shuaifei Chen, Ziheng Liu, Huahua Xiao, Derrick Wing Kwan Ng, Bo Ai

arXiv 2610.08387首次发表:更新:

发表机构

State Key Laboratory of Advanced Rail Autonomous Operation; School of Electronic and Information Engineering, Beijing Jiaotong University; School of Communications and Information Engineering, Xi’an University of Posts and Telecommunications; Purple Mountain Laboratories; ZTE Corporation; State Key Laboratory of Mobile Network and Mobile Multimedia Technology; University of New South Wales(先进轨道交通运行安全国家实验室; 北京交通大学电子与信息技术工程学院; 西安邮电大学通信与信息工程学院; 紫金山实验室; 中兴通讯股份有限公司; 移动网络与移动多媒体技术国家重点实验室; 新南威尔士大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对可移动天线辅助无蜂窝大规模MIMO系统,提出GLIIR-HAPPO多智能体强化学习算法,通过分解耦合优化、惩罚约束与图评论家协调,实现显著速率提升并接近集中式性能。

AI 中文摘要

可移动天线引入的固有非凸最小间距约束,对天线位置与传输策略的联合优化构成了严峻挑战,使得传统方法在计算上不可行,尤其是在大规模无蜂窝大规模多输入多输出(MIMO)系统中。在本工作中,我们提出了基于图学习的个体内在奖励异质智能体近端策略优化(GLIIR-HAPPO)算法,这是一种新颖的异质多智能体强化学习(MARL)框架,通过系统地将原始耦合优化问题分解为协调的子问题,从根本上克服了这一困境。为确保可解性,我们将非凸几何约束嵌入到惩罚增强的奖励结构中,并开发了专门的几何求解器,使定位智能体能够高效地在高维动作空间中导航。具体而言,我们提出了一种架构,包含用于自适应跨角色协调的动态交互图评论家,以及角色条件联邦蒸馏,通过紧凑的智能体输出统计量来同步同角色策略。在架构设计之外,我们建立了严格的理论分析,推导了单调性能改进界,并为所提出的双层优化建立了收敛保证。数值模拟表明,我们的框架相较于最先进的MARL方案实现了显著的求和速率提升。此外,我们先进架构的性能接近其完全集中式对应方案,同时大幅降低了通信开销。

英文摘要

The inherent non-convex minimum-separation constraints introduced by movable antennas present a formidable challenge to the joint optimization of antenna positions and transmission strategies, rendering conventional methods computationally infeasible, particularly in large-scale cell-free massive multiple-input multiple-output (MIMO). In this work, we propose the graph-based learning individual intrinsic reward heterogeneous-agent proximal policy optimization (GLIIR-HAPPO) algorithm, a novel heterogeneous multi-agent reinforcement learning (MARL) framework that fundamentally overcomes this impasse by systematically decomposing the original coupled optimization into coordinated subproblems. To ensure tractability, we embed the non-convex geometric constraints into a penalty-augmented reward structure and develop a specialized geometric solver that enables the positioning agents to efficiently navigate the high-dimensional action space. Specifically, we propose an architecture featuring a dynamic-interaction graph critic for adaptive cross-role coordination, together with role-conditioned federated distillation that synchronizes same-role policies through compact actor-output statistics. Beyond architectural design, we establish a rigorous theoretical analysis that derives monotonic performance improvement bounds and establishes convergence guarantees for the proposed bi-level optimization. Numerical simulations demonstrate that our framework yields significant sum-rate improvements over state-of-the-art MARL schemes. Moreover, the performance of our advanced architecture closely approaches its fully centralized counterpart, while drastically reducing communication overhead.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑