arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.14878cs.RO

基于MPC脚手架的真实世界强化学习用于灵巧操作

Real-World Reinforcement Learning with MPC Scaffolding for Dexterous Manipulation

发表机构本田美国研究院 · 佐治亚理工学院 · Skild AI
查看机构详情
  • Honda Research Institute USA(本田美国研究院)
  • Georgia Institute of Technology(佐治亚理工学院)
  • Skild AI

机构由 AI 辅助整理,请以论文原文为准。

Emek Barış Küçüktabak, Karankumar Patel, Zhaodong Yang, Jinda Cui, Kazuhiro Sasabuchi, Jun Takamatsu

首次发表
浏览论文内容

中文总结 AI 辅助

提出以MPC为脚手架的真实世界强化学习框架,通过预训练和在线引导加速灵巧操作学习,在16自由度手上7分钟达100%成功率,旋转速度超MPC五倍。

中文摘要 AI 辅助

真实世界强化学习(RL)为灵巧操作策略提供了一条有前景的途径,该策略可以直接从物理交互中适应,但学习过程受到早期探索效率低下和代价高昂的失败的阻碍。我们提出了一种框架,使用基于采样的模型预测控制(MPC)作为真实世界灵巧RL的脚手架,在学习过程中提供结构化的先验经验和任务导向的指导,无需人工演示或纠正动作。首先使用一小部分MPC轨迹填充离线回放缓冲区,并预训练演员和评论家网络。在线学习期间,MPC间歇性地指导数据收集,同时离策略的Soft Actor-Critic学习器从先前的MPC经验和新收集的物理交互中训练,控制逐渐过渡到学习到的策略。在16自由度Allegro手上进行连续的手内旋转任务中,该方法在7分钟的在线RL后达到了100%的策略单独评估成功率(5/5次试验),此前在硬件上收集了20条MPC轨迹,耗时12分钟。在线训练平均大约发生三次物体掉落。经过20分钟的在线学习,该策略的旋转速度是MPC控制器的五倍以上。它完成了超过110分钟内连续1000次旋转而无掉落。消融实验显示了基于MPC的预训练、保留的MPC经验和在线MPC指导的互补优势。我们进一步展示了对不同物体几何形状的快速适应以及成功的目标条件重定向,表明该框架能够实现高效、低干预的真实世界灵巧RL。

英文摘要

Real-world reinforcement learning (RL) offers a promising route to dexterous manipulation policies that can adapt directly from physical interaction, but learning is hindered by inefficient early exploration and costly failures. We propose a framework that uses sampling-based model predictive control (MPC) as scaffolding for real-world dexterous RL, providing structured prior experience and task-directed guidance during learning without human demonstrations or corrective actions. A small set of MPC trajectories is first used to populate an offline replay buffer and to pretrain the actor and critic. During online learning, MPC intermittently guides data collection while an off-policy Soft Actor-Critic learner trains from both prior MPC experience and newly collected physical interaction, with control gradually transitioning to the learned policy. On continuous in-hand rotation with a 16-DoF Allegro hand, initialized from 20 MPC trajectories collected in 12 minutes on hardware, the policy reaches 100\% success after 7 minutes of online RL, with about three object drops on average during training. After 20 minutes of online learning, the policy achieves more than five times the rotation speed of the MPC controller and completes 1000 consecutive rotations without a drop. Ablations show complementary benefits from MPC-based pretraining, retained MPC experience, and online MPC guidance, while additional experiments demonstrate rapid adaptation to new object geometries and successful goal-conditioned reorientation.

↑