arXivDaily arXiv每日学术速递 周一至周五更新

作者

Richard S. Sutton

Reinforcement Learning

2026-06-02 至 2026-06-02 共收录 1
2409.03915 2026-06-02 cs.LG math.OC

Asynchronous Stochastic Approximation with Applications to Average-Reward Reinforcement Learning

异步随机逼近及其在平均奖励强化学习中的应用

Huizhen Yu, Yi Wan, Richard S. Sutton

机构 * Department of Computing Science, University of Alberta(计算科学系,阿尔伯塔大学) Alberta Machine Intelligence Institute (Amii)(阿尔伯塔人工智能研究所(Amii))

AI总结 研究异步随机逼近算法的稳定性与收敛性,通过扩展Borkar-Meyn稳定性证明方法和Hirsch-Benaïm动力学系统方法,为平均奖励强化学习中的相对值迭代算法提供理论基础。

Comments 34 pages. This version contains only the asynchronous stochastic approximation material from version 2 of the original report; the reinforcement-learning material has been moved to a separate, stand-alone paper (arXiv:2512.06218). Minor corrections and additional remarks have been incorporated. A shorter version of this paper is to appear in the SIAM Journal on Control and Optimization

Journal ref SIAM Journal on Control and Optimization, 64(3):1456-1481, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏