arXivDaily arXiv每日学术速递 周一至周五更新

作者

Sergey Levine

Robotics / Reinforcement Learning

2026-06-04 至 2026-06-04 共收录 6
1506.02438 2026-06-04 cs.LG cs.RO cs.SY eess.SY

High-Dimensional Continuous Control Using Generalized Advantage Estimation

利用广义优势估计进行高维连续控制

John Schulman, Philipp Moritz, Sergey Levine, Michael Jordan, Pieter Abbeel

机构 * Department of Electrical Engineering and Computer Science(电气工程与计算机科学系) University of California, Berkeley(加州大学伯克利分校)

AI总结 本文提出了一种基于广义优势估计的方法,通过减少策略梯度估计的方差来解决高维连续控制中的样本需求问题,并通过信任区域优化提高稳定性和收敛性,从而在复杂的3D运动任务中实现了高效的政策学习。

详情

展开后加载摘要…

URL PDF HTML 收藏
1709.03153 2026-06-04 cs.LG cs.AI cs.RO cs.SY eess.SY

MBMF: Model-Based Priors for Model-Free Reinforcement Learning

MBMF:基于模型的先验用于无模型强化学习

Somil Bansal, Roberto Calandra, Kurtland Chua, Sergey Levine, Claire Tomlin

AI总结 本文提出一种结合模型与无模型强化学习的方法,通过学习概率动力学模型作为先验,提升数据效率和成本效益。

Comments After we submitted the paper for consideration in CoRL 2017 we found a paper published in the recent past with a similar method (see related work for a discussion). Considering the similarities between the two papers, we have decided to retract our paper from CoRL 2017

详情

展开后加载摘要…

URL PDF HTML 收藏
1703.09260 2026-06-04 eess.SY cs.LG cs.SY

Goal-Driven Dynamics Learning via Bayesian Optimization

通过贝叶斯优化的目标驱动动力学学习

Somil Bansal, Roberto Calandra, Ted Xiao, Sergey Levine, Claire J. Tomlin

AI总结 本文提出通过贝叶斯优化主动学习框架,迭代学习局部线性动力学模型以提升控制性能,用于四旋翼无人机任务控制。

Comments This is the extended version of the CDC'17 paper titled "Goal-Driven Dynamics Learning via Bayesian Optimization."

详情

展开后加载摘要…

URL PDF HTML 收藏
1611.05095 2026-06-04 cs.LG cs.RO cs.SY eess.SY

Learning Dexterous Manipulation Policies from Experience and Imitation

从经验与模仿中学习灵巧操作策略

Vikash Kumar, Abhishek Gupta, Emanuel Todorov, Sergey Levine

AI总结 本文研究了通过经验与模仿学习反馈控制灵巧五指手非抓取操作的任务,提出基于轨迹优化的局部控制器,并通过深度学习和最近邻方法进行泛化,展示了小数据训练下的有效性和盲操作优势。

Comments Initial draft for a journal submission

详情

展开后加载摘要…

URL PDF HTML 收藏
1603.00748 2026-06-04 cs.LG cs.AI cs.RO cs.SY eess.SY

Continuous Deep Q-Learning with Model-based Acceleration

基于模型的连续深度Q学习加速

Shixiang Gu, Timothy Lillicrap, Ilya Sutskever, Sergey Levine

AI总结 本文提出连续深度Q学习算法NAF及基于模型的加速方法,用于提升连续控制任务的样本效率和学习速度。

详情

展开后加载摘要…

URL PDF HTML 收藏
1311.1761 2026-06-04 cs.LG cs.AI cs.NE cs.RO cs.SY eess.SY

Exploring Deep and Recurrent Architectures for Optimal Control

探索深度和循环架构以实现最优控制

Sergey Levine

AI总结 本文探讨了将深度和循环神经网络应用于连续高维运动控制任务,通过强化学习算法训练控制器,比较不同架构的性能,并讨论深度学习在最优控制中的应用前景。

Comments Appears in the Neural Information Processing Systems (NIPS 2013) Workshop on Deep Learning

详情

展开后加载摘要…

URL PDF HTML 收藏