发表机构
Renmin University of China; Tsinghua University; University of Nottingham; Beijing Academy of Artificial Intelligence; Shanghai Qizhi Institute(中国人民大学; 清华大学; 诺丁汉大学; 北京人工智能研究院; 上海期智研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出ROBOCOACH,一个利用世界模型想象失败来指导演示获取和专家更新的教练框架,通过RIDI循环显著提升组合技能成功率,并在真实机器人上验证其有效性。
AI 中文摘要
长时程机器人操作在众多任务组合中复用技能,但通过额外的端到端演示来改进这些组合成本高昂。一个实用的自我改进系统必须决定接下来教什么以及在哪里应用这种监督。我们提出了ROBOCOACH,一个世界模型引导的教练框架,利用想象失败来指导演示请求和专家更新。其路线-想象-诊断-改进(RIDI)循环在COACHWORLD(我们共享的动作条件世界模型)中执行可复用的技能专家,并使用进度评判器记录第一个未能完成的子任务。聚合记录选择获取哪些子任务演示以及更新哪些专家适配器。在两个模拟套件和两个真实机器人平台上,想象与部署成功率在22个任务-策略对上相关(rho = 0.840)。对照比较表明,在匹配的数据预算和更新计划下,我们的教练方法优于匹配的基线。仅用150个额外的子任务演示,Franka上的成功率从13.3%提升至75.0%,AgileX上的成功率从40.0%提升至83.8%。被教练的专家还迁移到四个保留组合,平均成功率达到35.0%,而使用均匀获取演示的共享策略基线为0%。这些结果表明,世界模型可以作为主动教练,将想象失败转化为模块化策略改进的针对性监督。项目页面:this https URL
英文摘要
Long-horizon robot manipulation reuses skills across many task compositions, but improving these compositions with additional end-to-end demonstrations is costly. A practical self-improving system must decide both what to teach next and where to apply that supervision. We present ROBOCOACH, a world-model-guided coaching framework that uses imagined failures to guide demonstration requests and expert updates. Its Route-Imagine-Diagnose-Improve (RIDI) loop executes reusable skill experts inside COACHWORLD, our shared action-conditioned world model, and uses a progress judge to record the first subtask that fails to complete. Aggregated records select which subtask demonstrations to acquire and which expert adapters to update. Across two simulation suites and two real-robot platforms, imagined and deployed success correlate over 22 task-policy pairs (rho = 0.840). Controlled comparisons show that our coaching method outperforms matched baselines under matched data budgets and update schedules. With only 150 additional subtask demonstrations, success rises from 13.3% to 75.0% on Franka and from 40.0% to 83.8% on AgileX. The coached experts also transfer to four held-out compositions, achieving an average success of 35.0%, compared with 0% for a shared-policy baseline updated with uniformly acquired demonstrations. Together, these results show that world models can serve as active coaches, turning imagined failures into targeted supervision for modular policy improvement. Project Page: https://robocoach-ai.github.io/
Commentshttps://robocoach-ai.github.io/