arXivDaily arXiv每日学术速递 周一至周五更新

作者

Richard S. Sutton

Reinforcement Learning

共收录 77 篇
2608.01475 2026-08-04 cs.LG 新提交

Plasticity of Growing and Elastic Neural Networks in Online Continual Learning

在线持续学习中生长型与弹性型神经网络的可塑性

Jeong Min Kong, Richard S. Sutton

机构 * University of California, Los Angeles (UCLA)(加利福尼亚大学洛杉矶分校) ; University of Alberta(阿尔伯塔大学) ; Alberta Machine Intelligence Institute (Amii)(阿尔伯塔机器智能研究所)

AI总结 本文研究在线持续学习中生长型与弹性型神经网络的可塑性,实验显示自适应生长型和弹性型神经网络可维持高准确率、不丧失可塑性,或为在线持续学习提供有前景的算法类别。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.19357 2026-06-19 cs.RO cs.AI 新提交

Physical Atari: A Robust and Accessible Platform for Real-time Reinforcement Learning on Robots

Physical Atari: 一个用于机器人实时强化学习的鲁棒且可访问的平台

Khurram Javed, Joseph Modayil, Gloria Kennickell, Richard S. Sutton, John Carmack

机构 * Keen Technologies ; University of Alberta, Canada(阿尔伯塔大学,加拿大) ; Openmind Research Institute(Openmind研究机构)

AI总结 提出Physical Atari平台,通过机器人操作Atari控制器和实时渲染游戏帧,实现物理世界中的强化学习研究,验证了算法可直接在机器人上学习,并指出分布偏移会显著降低策略性能。

Comments To appear at RLC 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
1206.3285 2026-06-03 cs.AI cs.LG cs.SY eess.SY

Dyna-Style Planning with Linear Function Approximation and Prioritized Sweeping

具有线性函数逼近和优先级扫描的Dyna风格规划

Richard S. Sutton, Csaba Szepesvari, Alborz Geramifard, Michael P. Bowling

AI总结 本文提出一种基于模型的Dyna风格规划方法,扩展至线性函数逼近,证明其收敛性,并引入线性Dyna的优先级扫描算法。

Comments Appears in Proceedings of the Twenty-Fourth Conference on Uncertainty in Artificial Intelligence (UAI2008)

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.03915 2026-06-02 cs.LG math.OC

Asynchronous Stochastic Approximation with Applications to Average-Reward Reinforcement Learning

异步随机逼近及其在平均奖励强化学习中的应用

Huizhen Yu, Yi Wan, Richard S. Sutton

机构 * Department of Computing Science, University of Alberta(计算科学系,阿尔伯塔大学) ; Alberta Machine Intelligence Institute (Amii)(阿尔伯塔人工智能研究所(Amii))

AI总结 研究异步随机逼近算法的稳定性与收敛性,通过扩展Borkar-Meyn稳定性证明方法和Hirsch-Benaïm动力学系统方法,为平均奖励强化学习中的相对值迭代算法提供理论基础。

Comments 34 pages. This version contains only the asynchronous stochastic approximation material from version 2 of the original report; the reinforcement-learning material has been moved to a separate, stand-alone paper (arXiv:2512.06218). Minor corrections and additional remarks have been incorporated. A shorter version of this paper is to appear in the SIAM Journal on Control and Optimization

Journal ref SIAM Journal on Control and Optimization, 64(3):1456-1481, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.24238 2026-05-26 cs.AI

Toward Enactive Artificial Intelligence

走向生成式人工智能

Banafsheh Rafiee, Richard Sutton

机构 * Independent Researcher(独立研究者) ; Department of Computing Science, University of Alberta, Canada(阿尔伯塔大学计算机科学系) ; Alberta Machine Intelligence Institute (Amii)(阿尔伯塔机器智能研究所)

AI总结 本文主张将生成式认知方法融入人工智能,强调感知与行动不可分割、具身性和自主性,并指出强化学习在结构上与生成式原则存在共鸣但仍有差距。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.19033 2026-04-22 cs.LG cs.AI

Intentional Updates for Streaming Reinforcement Learning

为流式强化学习设计的意图更新

Arsalan Sharifnassab, Mohamed Elsayed, Kris De Asis, A. Rupam Mahmood, Richard S. Sutton

机构 * Openmind Research Institute(Openmind研究 institutes) ; Department of Computing Science, University of Alberta(阿尔伯塔大学计算机科学系)

AI总结 本文提出意图更新方法,通过定义更新目标来优化流式强化学习的稳定性与性能,实验表明其在流式环境下表现优异。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.06218 2025-12-09 cs.LG math.OC

Average-reward reinforcement learning in semi-Markov decision processes via relative value iteration

在半马尔可夫决策过程中的平均奖励强化学习中通过相对价值迭代

Huizhen Yu, Yi Wan, Richard S. Sutton

机构 * Department of Computing Science, University of Alberta(计算科学系,阿尔伯塔大学) ; Alberta Machine Intelligence Institute (Amii)(阿尔伯塔机器智能研究所)

AI总结 本文提出了一种在半马尔可夫决策过程中利用相对价值迭代算法进行平均奖励强化学习的方法,并引入新的单调性条件以提高算法收敛性。

Comments 24 pages. This paper presents the reinforcement-learning material previously contained in version 2 of arXiv:2409.03915, which is now being split into two stand-alone papers. Minor corrections and improvements to the main results have also been made in the course of this reformatting

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.19539 2025-07-29 cs.LG cs.AI stat.ML

Swift-Sarsa: Fast and Robust Linear Control

Swift-Sarsa:快速且鲁棒的线性控制

Khurram Javed, Richard S. Sutton

机构 * University of Alberta(阿尔伯塔大学)

AI总结 本文将SwiftTD的自适应步长机制与True Online Sarsa(λ)结合,提出线性同策略控制算法Swift-Sarsa,并在操作性条件反射基准上验证其能从海量含噪信号中识别相关特征并鲁棒分配信用。

Comments Presented at RLDM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.02342 2025-07-10 cs.LG cs.AI math.OC

MetaOptimize: A Framework for Optimizing Step Sizes and Other Meta-parameters

MetaOptimize:一种优化步长及其他元参数的框架

Arsalan Sharifnassab, Saber Salehkaleybar, Richard Sutton

机构 * Openmind Research Institute, Canada(开放心智研究机构,加拿大) ; Leiden Institute of Advanced Computer Science, Leiden University, Leiden, Netherlands(莱顿先进计算机科学研究所,莱顿大学,莱顿,荷兰) ; Department of Computing Science, University of Alberta, Edmonton, Canada(计算科学系,阿尔伯塔大学,爱德蒙顿,加拿大)

AI总结 MetaOptimize提出一种动态调整学习率等元参数的框架,可包裹一阶优化算法,通过最小化考虑长期影响的遗憾来优化训练,其低复杂度变体性能媲美最佳手工调度。

Journal ref ICML 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.09999 2024-10-31 cs.LG cs.AI

Reward Centering

Reward Centering(奖励中心化)

Abhishek Naik, Yi Wan, Manan Tomar, Richard S. Sutton

机构 * University of Alberta(阿尔伯塔大学) ; Alberta Machine Intelligence Institute(阿尔伯塔机器智能研究所) ; Meta AI

AI总结 该研究提出奖励中心化方法,通过减去奖励经验均值优化持续性强化学习的折扣方法,在常用折扣因子下提升显著且不受奖励常数平移影响,还给出离策略场景的均值估计方案,可广泛适配各类强化学习算法。

Comments In Proceedings of RLC 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.14951 2024-09-04 cs.LG cs.AI

An Idiosyncrasy of Time-discretization in Reinforcement Learning

强化学习中时间离散化的一个特性

Kris De Asis, Richard S. Sutton

机构 * University of Alberta(阿尔伯塔大学) ; Department of Computing Science(计算机科学系)

AI总结 本文研究强化学习中时间离散化粒度对回报定义的影响,指出直接应用离散时间算法存在不一致性,并提出一种简单修改以对齐连续时间与离散时间回报定义。

Comments RLC 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.16262 2024-08-30 cs.LG math.OC

On Convergence of Average-Reward Q-Learning in Weakly Communicating Markov Decision Processes

弱连通马尔可夫决策过程中平均报酬Q-学习的收敛性

Yi Wan, Huizhen Yu, Richard S. Sutton

机构 * University of Alberta(阿尔伯塔大学) ; Meta AI ; Alberta Machine Intelligence Institute (Amii)(阿尔伯塔机器智能研究所(Amii))

AI总结 本文将基于相对值迭代的Q-learning算法的几乎必然收敛性分析从单链扩展到弱连通MDPs,刻画了其收敛集性质,并推广至基于options框架的分层RL算法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.15091 2024-08-15 cs.LG math.OC

A Note on Stability in Asynchronous Stochastic Approximation without Communication Delays

Huizhen Yu, Yi Wan, Richard S. Sutton

机构 * University of Alberta(阿尔伯塔大学) ; Meta AI ; Alberta Machine Intelligence Institute (Amii)(阿尔伯塔机器智能研究所)

Comments Corrected typos and a minor error; parts of this material will be included in a separate future arXiv preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.14361 2024-07-23 cs.LG cs.AI

Auxiliary task discovery through generate-and-test

Banafsheh Rafiee, Sina Ghiassian, Jun Jin, Richard Sutton, Jun Luo, Adam White

机构 * University of Alberta(阿尔伯塔大学) ; Huawei Technologies Canada, Ltd(华为技术加拿大有限公司)

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.13812 2024-04-11 cs.LG

Maintaining Plasticity in Deep Continual Learning

Shibhansh Dohare, J. Fernando Hernandez-Garcia, Parash Rahman, A. Rupam Mahmood, Richard S. Sutton

机构 * University of Alberta(阿尔伯塔大学) ; Alberta Machine Intelligence Institute(阿尔伯塔机器智能研究所) ; CIFAR(加拿大高级研究所)

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.17401 2024-02-01 cs.LG cs.AI

Step-size Optimization for Continual Learning

Thomas Degris, Khurram Javed, Arsalan Sharifnassab, Yuxin Liu, Richard Sutton

机构 * DeepMind ; University of Alberta(阿尔伯塔大学)

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.01569 2023-12-27 cs.AI cs.LG

Iterative Option Discovery for Planning, by Planning

Kenny Young, Richard S. Sutton

机构 * University of Alberta(阿尔伯塔大学) ; Alberta Machine Intelligence Institute(阿尔伯塔机器智能研究所)

Comments Fixed incorrect arrows on some figures in the appendix

详情

展开后加载摘要…

URL PDF HTML 收藏
2202.03466 2023-09-19 cs.LG cs.AI

Reward-Respecting Subtasks for Model-Based Reinforcement Learning

Richard S. Sutton, Marlos C. Machado, G. Zacharias Holland, David Szepesvari, Finbarr Timbers, Brian Tanner, Adam White

机构 * DeepMind(深度思维) ; University of Alberta(阿尔伯塔大学) ; Alberta Machine Intelligence Institute (Amii)(阿尔伯塔机器智能研究所(Amii))

Journal ref Artificial Intelligence, first published online September 6, 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2301.13757 2023-07-25 cs.LG cs.AI

Toward Efficient Gradient-Based Value Estimation

Arsalan Sharifnassab, Richard Sutton

机构 * University of Alberta(阿尔伯塔大学)

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.15625 2023-06-28 cs.LG cs.AI

Value-aware Importance Weighting for Off-policy Reinforcement Learning

Kristopher De Asis, Eric Graves, Richard S. Sutton

机构 * University of Alberta(阿尔伯塔大学)

Comments CoLLAs 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2208.11173 2023-03-23 cs.AI cs.LG

The Alberta Plan for AI Research

Richard S. Sutton, Michael Bowling, Patrick M. Pilarski

机构 * University of Alberta(阿尔伯塔大学) ; Alberta Machine Intelligence Institute(阿尔伯塔机器智能研究所) ; DeepMind Alberta(DeepMind阿尔伯塔分部)

详情

展开后加载摘要…

URL PDF HTML 收藏
2207.01613 2022-11-29 cs.LG

Doubly-Asynchronous Value Iteration: Making Value Iteration Asynchronous in Actions

Tian Tian, Kenny Young, Richard S. Sutton

机构 * University of Alberta(阿尔伯塔大学) ; Alberta Machine Intelligence Institute(阿尔伯塔机器智能研究所)

详情

展开后加载摘要…

URL PDF HTML 收藏
2209.15141 2022-11-08 cs.LG

On Convergence of Average-Reward Off-Policy Control Algorithms in Weakly Communicating MDPs

Yi Wan, Richard S. Sutton

机构 * University of Alberta(阿尔伯塔大学) ; DeepMind(深度思维)

详情

展开后加载摘要…

URL PDF HTML 收藏
2101.02808 2022-10-19 cs.LG cs.AI

Average-Reward Off-Policy Policy Evaluation with Function Approximation

Shangtong Zhang, Yi Wan, Richard S. Sutton, Shimon Whiteson

机构 * University of Oxford(牛津大学) ; University of Alberta(阿尔伯塔大学)

Comments ICML 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2011.04590 2022-10-12 cs.AI

From Eye-blinks to State Construction: Diagnostic Benchmarks for Online Representation Learning

Banafsheh Rafiee, Zaheer Abbas, Sina Ghiassian, Raksha Kumaraswamy, Richard Sutton, Elliot Ludvig, Adam White

机构 * University of Alberta(阿尔伯塔大学) ; Alberta Machine Intelligence Institute (Amii)(阿尔伯塔机器智能研究所) ; DeepMind Alberta(DeepMind阿尔伯塔) ; University of Warwick(华威大学)

详情

展开后加载摘要…

URL PDF HTML 收藏
2205.12515 2022-10-03 cs.LG cs.AI

Toward Discovering Options that Achieve Faster Planning

Yi Wan, Richard S. Sutton

机构 * University of Alberta(阿尔伯塔大学) ; DeepMind(深度思维)

详情

展开后加载摘要…

URL PDF HTML 收藏
2202.13252 2022-06-07 cs.AI

The Quest for a Common Model of the Intelligent Decision Maker

Richard S. Sutton

机构 * DeepMind(深度思维) ; University of Alberta(阿尔伯塔大学)

Comments Will appear as an extended abstract at the fifth Multi-disciplinary Conference on Reinforcement Learning and Decision Making, held in Providence, Rhode Island, June 8-11, 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2108.06325 2022-05-06 cs.LG

Continual Backprop: Stochastic Gradient Descent with Persistent Randomness

Shibhansh Dohare, Richard S. Sutton, A. Rupam Mahmood

机构 * University of Alberta(阿尔伯塔大学) ; Alberta Machine Intelligence Institute (Amii)(阿尔伯塔机器智能研究所) ; Deepmind(深度思维)

详情

展开后加载摘要…

URL PDF HTML 收藏
2202.09701 2022-02-22 cs.LG

A History of Meta-gradient: Gradient Methods for Meta-learning

Richard S. Sutton

Comments 3 pages of text, 54 references

详情

展开后加载摘要…

URL PDF HTML 收藏
2112.15236 2022-01-03 cs.LG cs.AI

Learning Agent State Online with Recurrent Generate-and-Test

Amir Samani, Richard S. Sutton

机构 * University of Alberta(阿尔伯塔大学)

详情

展开后加载摘要…

URL PDF HTML 收藏
↑