arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

RoboRSI:复杂真实环境中稳定、高效且可复用的机器人自进化系统

RoboRSI: Stable, efficient, and reusable robot self-evolution in complex real-world environments

Zimo Wen, Yijin Chen, Yuxuan Cao, Wendi Chen, Yanwen Zou, Wenye Yu, Fuhang Kuang, Han Xue, Jun Lv, Chuan Wen, Cewu Lu

arXiv 2610.12424首次发表:更新:

AI 中文总结

RoboRSI是基于TSR的机器人自进化系统,可将任务分解为不同层级技能并整合稳定技能为可复用复合技能,在多目标家庭清洁任务及多个仿真基准上表现优于基线。

AI 中文摘要

通用机器人不仅应能执行多样化任务,还应能通过经验提升自身,将执行过程中习得的内容转化为后续任务可复用的能力。通过代码行动的机器人智能体已能根据执行反馈修复程序,但围绕赋予经验意义的任务结构来组织经验仍是核心挑战,需确保每次修复归因于负责的能力、有执行证据支持且在复用前经过验证。我们提出RoboRSI,这是一个基于自上而下技能细化(Top-Down Skill Refinement,TSR)构建的机器人自我改进系统。TSR将任务分解为具有明确职责范围和显式输入输出契约的复合技能、原子技能与基础技能,将每次执行结果归因于负责的分支,并将修订限制在该分支内。基于此结构,管理器、规划器、工程师与审核员协同完成规划、执行、诊断及新技能的验证发布,同时人类通过目标与修正引导该过程;稳定的技能序列会被进一步整合为可复用的复合技能。在移动机械臂上,RoboRSI在104轮迭代中开发了多目标家庭清洁任务;在仿真环境中,其在LIBERO、LIBERO-PRO、LIBERO-Plus及RoboTwin数据集上达到最高成功率,超出最强基线2.7至11.0个百分点。

英文摘要

A generalist robot should not only perform diverse tasks but also improve through experience, turning what it learns during execution into capabilities that later tasks can reuse. Robot agents that act through code can already repair programs from execution feedback, yet it remains a central challenge to organize this experience around the task structure that gives it meaning, so that each repair is attributed to the responsible capability, supported by execution evidence, and validated before it is reused. We introduce RoboRSI, a robot self-improvement system built on Top-Down Skill Refinement (TSR). TSR decomposes tasks into compound, atomic, and base skills with scoped responsibilities and explicit input--output contracts, attributes each execution outcome to the responsible branch, and confines revision to that branch. Building upon this structure, a Manager, Planner, Engineer, and Reviewer coordinate planning, execution, diagnosis, and the validated release of new skills, while people steer the process through objectives and corrections; stable skill sequences are further consolidated into reusable compound skills. On a mobile manipulator, RoboRSI develops multi-object household cleanup over 104 rounds. In simulation, it achieves the highest success rate on LIBERO, LIBERO-PRO, LIBERO-Plus, and RoboTwin, exceeding the strongest baseline by 2.7 to 11.0 percentage points.

Commentsproject page:https://lab.noematrix.ai/blog/2-roborsi/

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑