arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.02788cs.RO

Skill2Real:面向零样本仿真到现实机器人操作的智能体技能学习

Skill2Real: Agentic Skill Learning for Zero-Shot Sim-to-Real Robot Manipulation

Xincheng He, Siyu Ma, Chang Yu, Yunuo Chen, Yanjia Huang, Ying Nian Wu, Yin Yang, Chenfanfu Jiang

首次发表
浏览论文内容

中文总结 AI 辅助

Skill2Real提出基于API的智能体策略框架,通过PVG循环学习可迁移技能,在零样本下将LIBERO-Pro Long成功率从2.0%提升至56.3%,并实现真实世界78.75%的平均完成率。

中文摘要 AI 辅助

将机器人技能从仿真迁移到现实需要跨感知、动力学和具身差异仍可用的任务知识。我们提出Skill2Real,一种通过共享应用程序接口(API)学习可执行技能的智能体策略框架。提议者-验证者-治理者(PVG)循环利用特权仿真证据来诊断结果并验证更新,同时保持所学技能基于公共观察和API语义。小脑(Cerebellum)首先获取局部操作技能;大脑(Brain)随后在冻结小脑的情况下学习任务级组合。两种记忆均无需任务策略微调或技能记忆更新即可迁移到真实机器人。当GPT-5.6 Sol在LIBERO-90上学习技能时,用GPT-6 Astra评估每个冻结检查点将LIBERO-Pro Long成功率从2.0%提升至56.3%,而无需在Pro Long上训练。独立的Robosuite训练在七个任务上分别达到Sol和Opus 5的85.1%和89.4%平均成功率。冻结的Sol训练的LIBERO-90技能在四个真实世界操作任务中达到78.75%的平均完成率。在LIBERO-90训练期间移除验证者或治理者分别使最终Pro Long成功率降低17.3和13.3个百分点。这些结果支持通过通用机器人接口学习和迁移可执行技能的层次结构。

英文摘要

Transferring robotic skills from simulation to reality requires task knowledge that remains usable across differences in perception, dynamics, and embodiment. We introduce Skill2Real, an agentic policy framework that learns executable skills through a shared application programming interface (API). A Proposer-Verifier-Governor (PVG) loop uses privileged simulation evidence to diagnose outcomes and validate updates, while keeping learned skills grounded in public observations and API semantics. The Cerebellum first acquires local manipulation skills; the Brain then learns task-level composition with the Cerebellum frozen. Both memories transfer to the real robot without task-policy fine-tuning or skill-memory updates. As GPT-5.6 Sol learns skills on LIBERO-90, evaluating each frozen checkpoint with GPT-6 Astra raises LIBERO-Pro Long success from 2.0% to 56.3%, without training on Pro Long. Independent Robosuite training reaches 85.1% and 89.4% mean success with Sol and Opus 5 across seven tasks, respectively. Frozen Sol-trained LIBERO-90 skills achieve 78.75% mean completion across four real-world manipulation tasks with Astra. Removing the Verifier or Governor during LIBERO-90 training lowers final Pro Long success by 17.3 and 13.3 percentage points, respectively. These results support learning and transferring a hierarchy of executable skills through a common robot interface.

发表机构

  • University of California, Los Angeles(加州大学洛杉矶分校)
  • Shanghai Jiao Tong University(上海交通大学)
  • Envora
  • University of Utah(犹他大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑