arXivDaily arXiv每日学术速递 周一至周五更新

作者

Richard S. Sutton

Reinforcement Learning

2025-12-09 至 2025-12-09 共收录 1
2512.06218 2025-12-09 cs.LG math.OC

Average-reward reinforcement learning in semi-Markov decision processes via relative value iteration

在半马尔可夫决策过程中的平均奖励强化学习中通过相对价值迭代

Huizhen Yu, Yi Wan, Richard S. Sutton

机构 * Department of Computing Science, University of Alberta(计算科学系,阿尔伯塔大学) Alberta Machine Intelligence Institute (Amii)(阿尔伯塔机器智能研究所)

AI总结 本文提出了一种在半马尔可夫决策过程中利用相对价值迭代算法进行平均奖励强化学习的方法,并引入新的单调性条件以提高算法收敛性。

Comments 24 pages. This paper presents the reinforcement-learning material previously contained in version 2 of arXiv:2409.03915, which is now being split into two stand-alone papers. Minor corrections and improvements to the main results have also been made in the course of this reformatting

详情

展开后加载摘要…

URL PDF HTML 收藏