Average-reward reinforcement learning in semi-Markov decision processes via relative value iteration
在半马尔可夫决策过程中的平均奖励强化学习中通过相对价值迭代
机构 * Department of Computing Science, University of Alberta(计算科学系,阿尔伯塔大学) ; Alberta Machine Intelligence Institute (Amii)(阿尔伯塔机器智能研究所)
AI总结 本文提出了一种在半马尔可夫决策过程中利用相对价值迭代算法进行平均奖励强化学习的方法,并引入新的单调性条件以提高算法收敛性。
Comments 24 pages. This paper presents the reinforcement-learning material previously contained in version 2 of arXiv:2409.03915, which is now being split into two stand-alone papers. Minor corrections and improvements to the main results have also been made in the course of this reformatting