AI 中文总结
研究两跳中继网络中合作博弈的节能功率控制问题,在难以获取瞬时CSI的非理想情况下,提出基于延迟奖励的状态 - 动作值函数和多智能体深度Q网络学习框架,该方法优于其他方法且接近最优解。
AI 中文摘要
本文研究合作通信网络中的合作博弈,其中各中继自主决策,目标是实现相同的最大化能量效率优化目标。考虑难以获取瞬时信道状态信息(CSI),仅有部分可观测过时CSI的非理想情况。为解决此博弈问题,定义基于延迟奖励的状态 - 动作值函数并提出多智能体深度Q网络学习框架。通过分析证明基于瞬时CSI的博弈论方法所得效用为此方法效用的上限。仿真结果表明该方法显著优于其他方法,与最优解仅相差约5.2%。
英文摘要
In this paper, we study a cooperative game in the cooperative communication network, where each relay makes decisions autonomously and aims to achieve the same optimization objective of maximizing energy efficiency. We consider the non-ideal situation where instantaneous channel state information (CSI) is difficult to obtain and only partially observable outdated CSI is available. To solve this game problem, we define a delayed reward-based state-action value function and propose a multi-agent deep Q network learning framework. Then, we prove analytically that utilities obtained by game-theoretic approaches with the instantaneous CSI serve as upper bounds for those of the proposed method. Simulation results reveal that our approach considerably outperforms its potential alternatives and is only about 5.2% away from the optimal solution.