Asynchronous Stochastic Approximation with Applications to Average-Reward Reinforcement Learning
异步随机逼近及其在平均奖励强化学习中的应用
机构 * Department of Computing Science, University of Alberta(计算科学系,阿尔伯塔大学) ; Alberta Machine Intelligence Institute (Amii)(阿尔伯塔人工智能研究所(Amii))
AI总结 研究异步随机逼近算法的稳定性与收敛性,通过扩展Borkar-Meyn稳定性证明方法和Hirsch-Benaïm动力学系统方法,为平均奖励强化学习中的相对值迭代算法提供理论基础。
Comments 34 pages. This version contains only the asynchronous stochastic approximation material from version 2 of the original report; the reinforcement-learning material has been moved to a separate, stand-alone paper (arXiv:2512.06218). Minor corrections and additional remarks have been incorporated. A shorter version of this paper is to appear in the SIAM Journal on Control and Optimization
Journal ref SIAM Journal on Control and Optimization, 64(3):1456-1481, 2026