AI 中文总结
FedCritic-MIMO是面向开放解耦6G RAN的通信高效无服务器联邦多智能体强化学习框架,通过稀疏评论者交换等技术实现多小区大规模MIMO资源控制,性能最优且通信开销降低76%。
AI 中文摘要
本文提出FedCritic-MIMO,一种通信高效的无服务器联邦多智能体强化学习框架,用于开放与解耦6G无线接入网(RAN)中独立部署的小区级控制器的AI原生资源控制。各控制器无共享训练器,保留本地执行器与个性化评论者组件,仅交换兼容的共享评论者参数。FedCritic-MIMO针对复用因子为1的多小区大规模MIMO正交频分多址(OFDMA)部署场景,其中RAN控制器需在有限的控制器间信令下协同管理用户调度、每流功率分配、波束成形、干扰及长期服务质量(QoS)。每个基站在无集中式训练或执行器联邦的情况下本地执行其执行器,而评论者知识通过感知干扰的图以对等方式交换。该框架通过感知无线环境的事件触发、带误差反馈的自适应分层前k稀疏评论者交换,以及平衡感知干扰的融合实现协作。我们在固定策略、冻结目标评论者回归模型下,为平衡压缩的对等评论者递归建立了有限时间平稳性与一致性保证。在强干扰耦合的复用因子为1的仿真中,FedCritic-MIMO在启发式方法、独立学习、集中式训练及通信消融基线中实现了最优的性能-通信权衡;其在学习基线中取得最高的保留吞吐量,改善用户速率分布与平均信干噪比(SINR),提升QoS满意度,且每传输比特的干扰成本最低;与未压缩的分布式评论者交换相比,其评论者通信开销降低76%。这些结果表明,兼容共享评论者参数的无服务器交换可在无需集中式轨迹收集或参数服务器聚合的情况下协调RAN控制器。
英文摘要
This paper proposes FedCritic-MIMO, a communication-efficient serverless federated multi-agent reinforcement learning framework for AI-native resource control across independently deployable cell-level controllers in open and disaggregated 6G RANs. Controllers share no trainer, retain local actors and personalized critic components, and exchange only compatible shared critic parameters. FedCritic-MIMO targets reuse-$1$ multi-cell massive-MIMO OFDMA deployments, where RAN controllers jointly manage user scheduling, per-stream power allocation, beamforming, interference, and long-term QoS with limited inter-controller signaling. Each base station locally executes its actor without centralized training or actor federation, while critic knowledge is exchanged peer-to-peer over an interference-aware graph. It enables this collaboration through wireless-aware event triggering, adaptive layer-wise top-$k$ sparse critic exchange with error feedback, and balanced interference-aware fusion. We establish conditional finite-time stationarity and consensus guarantees for the balanced, compressed peer-to-peer critic recursion under a fixed-policy, frozen-target critic-regression model. In strongly interference-coupled reuse-$1$ simulations, FedCritic-MIMO achieves the best performance-communication tradeoff among heuristic, independent-learning, centralized-training, and communication-ablation baselines. It achieves the highest held-out throughput, improves user-rate distribution and mean SINR, increases QoS satisfaction, and attains the lowest interference cost per delivered bit among learning baselines. It reduces critic-communication overhead by $76\%$ relative to uncompressed distributed critic exchange. These results demonstrate that serverless exchange of compatible shared critic parameters can coordinate RAN controllers without centralized trajectory collection or parameter-server aggregation.
CommentsSubmitted to IEEE for possible publication