发表机构
UC Berkeley; MIT; Amazon FAR; University of Chicago(加州大学伯克利分校; 麻省理工学院; 亚马逊FAR; 芝加哥大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出RPG框架,通过离线数据识别技能、仿真练习与诊断改进,无需更新权重即可提升具身智能体任务成功率,从28.6%提升至95.0%,并完成30次物理试验。
AI 中文摘要
构建跨多样化任务的可靠机器人能力需要大量人力来开发和维护技能、设计奖励函数以及整合感知与控制。我们提出了重建、练习、实战(RPG)框架,该框架在不更新模型权重的情况下实现机器人执行系统的自主改进。RPG在离线数据集中识别操作能力,并在仿真中构建相关的练习任务。在练习过程中,RPG利用执行反馈、特权仿真器状态和可用的数据集视频来诊断失败原因。它基于这些诊断开发新的可复用符号技能、改进现有技能并修订系统提示。跨任务评估在保留重用之前测试各个候选更改和合并后的修订。在测试时,多模态大语言模型利用最终的系统提示和技能库来协调感知与机器人控制。在22个操作任务的保留初始化上,RPG将任务成功率从第一轮练习后的28.6%提升到15轮后的95.0%,优于所有评估的基线方法,包括ASPIRE(75.5%)和由GPT-6 Astra Pro驱动的CaP-Agent0(60.0%)。经过常见的校准和硬件适配流程后,冻结系统在全部30次物理试验中成功,其中三个任务各进行十次试验。项目网站:此https URL
英文摘要
Building reliable robot capabilities across diverse tasks requires substantial human effort to develop and maintain skills, design rewards, and integrate perception with control. We present Reconstruct, Practice, Go Real (RPG), a framework for autonomous improvement of robot execution systems without updating model weights. RPG identifies manipulation capabilities in an offline dataset and constructs related practice tasks in simulation. During practice, RPG uses execution feedback, privileged simulator state, and available dataset videos to diagnose failures. It develops new reusable symbolic skills, refines existing skills, and revises the system prompt based on these diagnoses. Cross-task evaluation tests individual candidate changes and merged revisions before they are retained for reuse. At test time, a multimodal LLM uses the resulting system prompt and skill library to coordinate perception and robot control. On held-out initializations of 22 manipulation tasks, RPG improves task success from 28.6% after the first practice round to 95.0% after 15 rounds, outperforming all evaluated baselines, including ASPIRE (75.5%) and CaP-Agent0 powered by GPT-6 Astra Pro (60.0%). After a common calibration and hardware-adaptation procedure, the frozen system succeeds in all 30 physical trials, with ten trials on each of three tasks. Project Website: https://rpg-robot.github.io/
Comments17 pages, 6 figures, 10 tables