arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

作者

Richard S. Sutton

Reinforcement Learning

至 收录 77
2110.13855 2021-10-27 cs.LG

Average-Reward Learning and Planning with Options

Yi Wan, Abhishek Naik, Richard S. Sutton

URL PDF HTML 收藏
2109.05110 2021-09-14 cs.LG cs.AI

An Empirical Comparison of Off-policy Prediction Learning Algorithms in the Four Rooms Environment

Sina Ghiassian, Richard S. Sutton

Comments 13 pages

URL PDF HTML 收藏
2006.16318 2021-06-29 cs.LG cs.AI

Learning and Planning in Average-Reward Markov Decision Processes

Yi Wan, Abhishek Naik, Richard S. Sutton

Comments In Proceedings of ICML 2021

URL PDF HTML 收藏
2106.00922 2021-06-15 cs.LG cs.AI

An Empirical Comparison of Off-policy Prediction Learning Algorithms on the Collision Task

Sina Ghiassian, Richard S. Sutton

URL PDF HTML 收藏
2102.07686 2021-06-10 cs.LG cs.AI stat.ML

Does the Adam Optimizer Exacerbate Catastrophic Forgetting?

Dylan R. Ashley, Sina Ghiassian, Richard S. Sutton

Comments 9 pages in main text + 3 pages of references + 16 pages of appendices, 6 figures in main text + 21 figures in appendices, 6 tables in appendices; source code available at https://github.com/dylanashley/catastrophic-forgetting/tree/arxiv

URL PDF HTML 收藏
2104.08543 2021-04-20 cs.AI

Planning with Expectation Models for Control

Katya Kudashkina, Yi Wan, Abhishek Naik, Richard S. Sutton

URL PDF HTML 收藏
1705.03520 2021-04-06 cs.AI cs.LG cs.SY eess.SY

Policy Iterations for Reinforcement Learning Problems in Continuous Time and Space -- Fundamental Theory and Methods

Jaeyoung Lee, Richard S. Sutton

Comments To appear in Automatica. All the Appendices are provided

Journal ref Automatica vol. 126, 109421 (2021)

URL PDF HTML 收藏
2010.15268 2020-10-30 cs.LG cs.AI

Understanding the Pathologies of Approximate Policy Evaluation when Combined with Greedification in Reinforcement Learning

Kenny Young, Richard S. Sutton

URL PDF HTML 收藏
2008.12095 2020-08-28 cs.AI cs.HC cs.LG

Document-editing Assistants and Model-based Reinforcement Learning as a Path to Conversational AI

Katya Kudashkina, Patrick M. Pilarski, Richard S. Sutton

Comments Currently under review

URL PDF HTML 收藏
2008.11329 2020-08-27 cs.LG cs.AI

Inverse Policy Evaluation for Value-based Sequential Decision-making

Alan Chan, Kris de Asis, Richard S. Sutton

Comments Submitted to NeurIPS 2020

URL PDF HTML 收藏
1904.01191 2020-07-31 cs.LG cs.AI stat.ML

Planning with Expectation Models

Yi Wan, Zaheer Abbas, Adam White, Martha White, Richard S. Sutton

URL PDF HTML 收藏
1908.03568 2020-02-17 cs.LG cs.AI stat.ML

Behaviour Suite for Reinforcement Learning

Ian Osband, Yotam Doron, Matteo Hessel, John Aslanides, Eren Sezener, Andre Saraiva, Katrina McKinney, Tor Lattimore, Csaba Szepesvari, Satinder Singh, Benjamin Van Roy, Richard Sutton, David Silver, Hado Van Hasselt

URL PDF HTML 收藏
1909.03906 2020-02-12 cs.LG cs.AI

Fixed-Horizon Temporal Difference Methods for Stable Reinforcement Learning

Kristopher De Asis, Alan Chan, Silviu Pitis, Richard S. Sutton, Daniel Graves

Comments AAAI 2020

URL PDF HTML 收藏
1912.04002 2019-12-10 cs.LG stat.ML

Learning Sparse Representations Incrementally in Deep Reinforcement Learning

J. Fernando Hernandez-Garcia, Richard S. Sutton

URL PDF HTML 收藏
1910.02140 2019-11-28 cs.AI

Discounted Reinforcement Learning Is Not an Optimization Problem

Abhishek Naik, Roshan Shariff, Niko Yasui, Hengshuai Yao, Richard S. Sutton

Comments Accepted for presentation at the Optimization Foundations of Reinforcement Learning Workshop at NeurIPS 2019

URL PDF HTML 收藏
1903.03252 2019-03-11 cs.LG cs.AI stat.ML

Learning Feature Relevance Through Step Size Adaptation in Temporal-Difference Learning

Alex Kearney, Vivek Veeriah, Jaden Travnik, Patrick M. Pilarski, Richard S. Sutton

URL PDF HTML 收藏
1903.00194 2019-03-04 cs.AI cs.LG

Should All Temporal Difference Learning Use Emphasis?

Xiang Gu, Sina Ghiassian, Richard S. Sutton

URL PDF HTML 收藏
1901.07510 2019-02-11 cs.LG stat.ML

Understanding Multi-Step Deep Reinforcement Learning: A Systematic Study of the DQN Target

J. Fernando Hernandez-Garcia, Richard S. Sutton

URL PDF HTML 收藏
1704.04463 2018-11-27 cs.LG math.OC

On Generalized Bellman Equations and Temporal-Difference Learning

Huizhen Yu, A. Rupam Mahmood, Richard S. Sutton

Comments Minor revision; 41 pages; to appear in Journal on Machine Learning Research, 2018

Journal ref Journal of Machine Learning Research 19(48):1-49, 2018

URL PDF HTML 收藏
1811.02597 2018-11-08 cs.LG cs.AI stat.ML

Online Off-policy Prediction

Sina Ghiassian, Andrew Patterson, Martha White, Richard S. Sutton, Adam White

Comments 68 pages

URL PDF HTML 收藏
1809.07435 2018-09-21 cs.LG cs.AI eess.SP

Predicting Periodicity with Temporal Difference Learning

Kristopher De Asis, Brendan Bennett, Richard S. Sutton

URL PDF HTML 收藏
1807.01830 2018-09-10 cs.LG cs.AI stat.ML

Per-decision Multi-step Temporal Difference Learning with Control Variates

Kristopher De Asis, Richard S. Sutton

Journal ref (2018). In Conference on Uncertainty in Artificial Intelligence. http://auai.org/uai2018/proceedings/papers/282.pdf

URL PDF HTML 收藏
1805.07476 2018-09-10 cs.LG cs.AI stat.ML

Two geometric input transformation methods for fast online reinforcement learning with neural nets

Sina Ghiassian, Huizhen Yu, Banafsheh Rafiee, Richard S. Sutton

Comments 16 pages

URL PDF HTML 收藏
1802.06139 2018-06-29 cs.AI cs.LG

Reactive Reinforcement Learning in Asynchronous Environments

Jaden B. Travnik, Kory W. Mathewson, Richard S. Sutton, Patrick M. Pilarski

Comments 11 pages, 7 figures, currently under journal peer review

URL PDF HTML 收藏
1703.01327 2018-06-13 cs.AI cs.LG

Multi-step Reinforcement Learning: A Unifying Algorithm

Kristopher De Asis, J. Fernando Hernandez-Garcia, G. Zacharias Holland, Richard S. Sutton

Comments Appeared at the Thirty-Second AAAI Conference on Artificial Intelligence (AAAI-18)

Journal ref (2018). In AAAI Conference on Artificial Intelligence. https://www.aaai.org/ocs/index.php/AAAI/AAAI18/paper/view/16294

URL PDF HTML 收藏
1806.00540 2018-06-05 cs.LG cs.AI stat.ML

Integrating Episodic Memory into a Reinforcement Learning Agent using Reservoir Sampling

Kenny J. Young, Richard S. Sutton, Shuo Yang

URL PDF HTML 收藏
1712.01275 2018-05-01 cs.LG cs.AI

A Deeper Look at Experience Replay

Shangtong Zhang, Richard S. Sutton

Comments NIPS 2017 Deep Reinforcement Learning Symposium

URL PDF HTML 收藏
1804.03334 2018-04-11 cs.LG stat.ML

TIDBD: Adapting Temporal-difference Step-sizes Through Stochastic Meta-descent

Alex Kearney, Vivek Veeriah, Jaden B. Travnik, Richard S. Sutton, Patrick M. Pilarski

Comments Version as submitted to the 31st Conference on Neural Information Processing Systems (NIPS 2017) on May 19, 2017. 9 pages, 5 figures. Extended version in preparation for journal submission

URL PDF HTML 收藏
1801.08287 2018-02-15 cs.AI

Directly Estimating the Variance of the λ-Return Using Temporal-Difference Methods

Craig Sherstan, Brendan Bennett, Kenny Young, Dylan R. Ashley, Adam White, Martha White, Richard S. Sutton

URL PDF HTML 收藏
1711.03676 2017-11-13 cs.AI cs.HC cs.LG

Communicative Capital for Prosthetic Agents

Patrick M. Pilarski, Richard S. Sutton, Kory W. Mathewson, Craig Sherstan, Adam S. R. Parker, Ann L. Edwards

Comments 33 pages, 10 figures; unpublished technical report undergoing peer review

URL PDF HTML 收藏