EmbodiedRSI:通过假设引导的协同进化实现主动持续机器人学习
EmbodiedRSI: Active Continual Robot Learning Through Hypothesis-Guided Co-Evolution
浏览论文内容
中文总结 AI 辅助
EmbodiedRSI通过快慢双系统架构和假设图,自主选择实验并协同进化代码与技能,在多个基准和真实机器人上显著提升持续学习性能。
中文摘要 AI 辅助
机器人基础模型提供了强大的视觉运动控制能力,但当物体位置或任务指令发生变化时,其性能可能会下降。进一步的改进通常需要在大量机器人数据上进行后训练,而通过遥操作等方法收集这些数据可能成本高昂。智能体化(Agentic)框架可以在模型周围进行适应,但当前的自进化框架在决定要追求哪些代码和技能变更时,对机器人试验的使用效率低下。我们引入了EmbodiedRSI,一个自进化的智能体化框架,它自主决定下一步探索哪里,并将由此产生的物理交互转化为改进的代码和技能。EmbodiedRSI通过一种快慢双系统架构(Fast-Slow Dual-System Architecture)实现这一点,其中相互竞争的代码和技能假设被维护在一个假设图(Hypothesis Graph)中。信息价值实验选择(Value-of-Information Experiment Selection)会选择能够区分这些假设的物理实验。这些实验的结果指导代码-技能协同进化(Code-Skill Co-Evolution)。慢系统构建层次记忆(Hierarchical Memory),而奖励基础记忆学习(Reward-Grounded Memory Learning)根据记忆对后续快系统改进的价值来选择有效的记忆。在RoboCasa365上,EmbodiedRSI达到了77.0%的总体成功率和71.3%的Composite-Unseen成功率,而最佳基线的成功率仅为40.1%。EmbodiedRSI在LIBERO-Pro上也达到了86.8%的总体成功率。除了基准性能外,EmbodiedRSI还零样本迁移到真实世界机器人,在多个具有挑战性的任务中实现了71.3%的总体成功率。
英文摘要
Robot foundation models provide strong visuomotor control, yet their performance can degrade when object positions or task instructions change. Further improvements often require post-training on substantial robot data, which can be costly to collect through methods such as teleoperation. Agentic harnesses can adapt around the model, but current self-evolving harnesses use robot trials inefficiently when deciding which code and skill changes to pursue. We introduce EmbodiedRSI, a self-evolving agentic harness that autonomously decides where to explore next and turns the resulting physical interaction into improved code and skills. EmbodiedRSI realizes this through a Fast-Slow Dual-System Architecture, in which competing code and skill hypotheses are maintained in a Hypothesis Graph. Value-of-Information Experiment Selection chooses physical experiments that can distinguish these hypotheses. Their outcomes guide Code-Skill Co-Evolution. The Slow System builds Hierarchical Memory, and Reward-Grounded Memory Learning selects effective memory according to their value for later Fast-System improvement. On RoboCasa365, EmbodiedRSI reaches 77.0% overall success and 71.3% on Composite-Unseen, compared with 40.1% for the best baseline. EmbodiedRSI also reaches 86.8% overall success on LIBERO-Pro. Beyond benchmark performance, EmbodiedRSI transfers zero-shot to real-world robot, achieving 71.3% overall success across multiple challenging tasks.
发表机构
- Columbia University(哥伦比亚大学)
- Princeton University(普林斯顿大学)
机构由 AI 辅助整理,请以论文原文为准。