arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

动态OWC网络中面向能效的协作式Dueling DQN SAC学习

Cooperative Dueling DQN SAC Learning for Energy Efficiency in Dynamic OWC Networks

Walter Zibusiso Ncube, Ahmad Adnan Qidan, Taisir El-Gorashi, Jaafar M. H. Elmirghani

arXiv 2610.09266首次发表:更新:

发表机构

King’s College London(伦敦国王学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对动态OWC网络的能效优化问题,提出DARA-DRL双智能体强化学习方法,结合Dueling DQN与SAC,实现联合资源分配,能效提升28.8%,QoS提升10.5%。

AI 中文摘要

日益增长的无线流量正加剧拥挤的射频频谱的压力。光无线通信(OWC)通过利用丰富的免许可光频谱提供了一种补充解决方案。然而,室内OWC网络是动态的:用户会移动、进入或离开网络,并订阅不同的服务。协调不当的资源分配因此可能浪费子载波、需要过高的传输功率并导致频繁的AP重新分配。能效(EE)定义为总传输数据速率除以总网络功耗,因此要求在维持服务质量(QoS)的同时,联合调整服务AP、分配的子载波数量和传输功率。在动态时间序列、多服务OWC环境中优化这些因素,产生了一个复杂的序列化能效优化问题。为解决该问题,本工作提出了基于深度强化学习的双智能体资源分配(DARA-DRL)。DARA-DRL结合了用于关联和子载波分配的分支Dueling深度Q网络,以及用于连续功率控制的条件软演员-评论家智能体。这两个智能体通过共同奖励和协作式价值更新耦合,该更新将每个离散分配与其相应的功率决策一起评估。仿真结果表明,DARA-DRL保持在最优解的5%以内,并且与最先进的基准相比,它将能效提高了28.8%,QoS满意度提高了10.5%,同时将在线决策时间减少了14.5%。结果证明,智能体专业化简化了混合动作学习,并且协作优于独立训练的智能体。

英文摘要

Growing wireless traffic is increasing pressure on the congested radio-frequency spectrum. Optical wireless communication (OWC) provides a complementary solution by using the abundant unlicensed optical spectrum. However, indoor OWC networks are dynamic: users move, enter or leave the network, and subscribe to different services. Poorly coordinated resource allocation can consequently waste subcarriers, require excessive transmission power and cause frequent AP reassignments. Energy efficiency (EE), defined as the total delivered data rate divided by the total network power consumption, therefore requires the serving AP, number of allocated subcarriers, and transmission power to be jointly adapted while maintaining QoS. Optimising these in a dynamic time series, multi-service OWC environment produces a complex sequential EE optimisation problem. To address this problem, this work proposes Dual-Agent Resource Allocation using Deep Reinforcement Learning (DARA-DRL). DARA-DRL combines a branching duelling deep Q-network for association and subcarrier allocation with a conditional soft actor-critic agent for continuous power control. The agents are coupled through a common reward and a cooperative value update that evaluates each discrete allocation together with its corresponding power decision. Simulation results show that DARA-DRL remains within 5\% of the optimal solution and, compared with state-of-the-art benchmarks, it improves EE by 28.8\% and QoS satisfaction by 10.5\%, while reducing online decision time by 14.5\%. Results demonstrate that agent specialisation simplifies mixed-action learning, and cooperation outperforms independently trained agents.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑