arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.19049cs.LGeess.SP

面向智慧校园覆盖的多智能体离线策略深度强化学习

Multi-Agent Off-Policy Deep Reinforcement Learning for Smart Campus Coverage

Omar Rady, Mohamed Ayman, Ali Arafa, Mohamed Shalma

AI总结:

本文针对现实非凸校园拓扑下的毫米波基站最优部署问题,将其建模为马尔可夫决策过程,通过对比四种深度强化学习方案,发现地理划分多智能体DDPG在密集场景中性能更优、可实现全覆盖且收敛高效。

AI中文摘要:

深度强化学习(DRL)因具备实时自适应能力及在复杂优化问题中的有效性,近来受到广泛关注。本文研究在现实非凸校园拓扑中部署毫米波(mmWave)基站(BS)的最优方案,该优化问题因最大最小公平性目标的非凸、非平滑特性属于NP难问题。为克服这些约束,我们将基站部署建模为马尔可夫决策过程(MDP),并系统评估四种DRL方案:离散单智能体深度Q网络(DQN)、空间划分多智能体DQN、连续单智能体深度确定性策略梯度(DDPG)、地理划分多智能体DDPG框架。数值评估显示,在密集场景下多智能体DDPG方法的性能显著优于单智能体,可实现全覆盖,公平性Jain指数达0.94,且在含400个用户的密集场景中展现出高效的计算收敛性。

英文摘要:

Deep reinforcement learning (DRL) has recently gained a great attention due to its real-time adaptation and effectiveness in complex optimization problems. This paper investigates the optimal deployment of millimeter-wave (mmWave) base stations (BSs) in a realistic, non-convex campus topology. The optimization problem is NP-hard, due to the non-convex, non-smooth nature of the max-min fairness objective. To overcome these constraints, we formulate the BS placement as a Markov Decision Process (MDP) and systematically benchmark four DRL schemes: a discrete single-agent Deep Q-Network (DQN), a spatially partitioned Multi-Agent DQN, a continuous single-agent Deep Deterministic Policy Gradient (DDPG), and a geographically partitioned multi-agent DDPG framework. Numerical evaluations reveal that the multi-agent DDPG approach substantially outperforms single-agent in dense scenarios. Additionally full coverage is achieved, and a fairness Jain's index of 0.94 is obtained. Finally, the multi-agent demonstrates highly efficient computational convergence of dense scenarios with $400$ users.

↑