arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.01651cs.RO

一个演示足以实现真实世界机器人强化学习

One Demonstration Is Enough for Real-World Robotic Reinforcement Learning

Yuwan Liu, Hongze Yu, Song Liu, Yuhan Wang, Junge Zhang, Yaodong Yang, Yuanpei Chen, Ceyao Zhang

首次发表
浏览论文内容

中文总结 AI 辅助

提出AutoSERL框架,仅需一个演示即可自动化真实世界机器人强化学习,通过滑动窗口干预、安全恢复和自动终止机制,在六个接触密集型任务上超越基线方法。

中文摘要 AI 辅助

在物理硬件上学习有效的机器人控制策略具有挑战性,原因在于数据收集成本高且奖励规范困难。先前的工作已将演示纳入强化学习(RL),但现有方法要么需要大量演示,要么在训练过程中依赖持续的人类干预。为了解决这些限制,我们提出了AutoSERL,一个利用单个演示完全自动化真实世界机器人RL中干预过程的框架。该框架包括三种互补机制以完成特定任务:滑动窗口干预机制,持续引导探索以防止局部最优和不安全偏离;安全恢复机制,通过预定义的轨迹恢复点检测并纠正失败状态;以及干预终止标准,一旦策略能够独立完成任务,自动禁用引导,保留其探索优势。我们在两个机器人平台的六个接触密集型操作任务上评估了AutoSERL,涵盖插入、悬挂和基于铰链的任务。AutoSERL在所有任务上一致优于使用20个演示初始化的SERL、行为克隆和MILES(一种专用的一次性模仿学习基线),同时匹配HIL-SERL,在插入任务上达到100%的成功率,并展示了对位置变化的更强鲁棒性,所有这些都仅来自单个演示。代码和视频可在我们的项目网站上获取:this https URL。

英文摘要

Learning effective robot control policies on physical hardware is challenging due to costly data collection and the difficulty of reward specification. Prior work has incorporated demonstrations into reinforcement learning (RL), yet existing approaches either require large numbers of demonstrations or depend on continuous human intervention during training. To address these limitations, we present AutoSERL, a framework that leverages a single demonstration to fully automate the intervention process in real-world robot RL. The framework includes three complementary mechanisms to accomplish certain tasks: a sliding window intervention mechanism that continuously guides exploration to prevent local optima and unsafe deviations, a safety recovery mechanism that detects and corrects failure states via predefined trajectory recovery points, and an intervention termination criterion that automatically disables guidance once the policy can independently complete the task, preserving its exploration advantage. We evaluate AutoSERL on six contact-intensive manipulation tasks across two robot platforms, spanning insertion, hanging, and hinge-based tasks. AutoSERL consistently outperforms SERL initialized with 20 demonstrations, behavior cloning, and MILES -- a dedicated one-shot imitation learning baseline -- across all tasks while matching HIL-SERL, achieves 100% success rate on insertion tasks, and demonstrates improved robustness to positional variations, all from a single demonstration. Code and videos are available on our project website: https://autoserl.github.io/.

发表机构

  • National Key Laboratory of Cognition and Decision Intelligence for Complex Systems, Institution of Automation, Chinese Academy of Sciences(中国科学院自动化研究所复杂系统认知与决策智能国家重点实验室)
  • Beijing Academy of Artificial Intelligence(北京智源人工智能研究院)
  • PKU-PsiBot Joint Lab(北大-鹏城实验室联合实验室)
  • School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院)
  • Institute for Artificial Intelligence, Peking University(北京大学人工智能研究院)

机构由 AI 辅助整理,请以论文原文为准。

↑