F4R:面向持续机器人自我改进的失败驱动识别、重建、精炼与重新部署
F4R: Failure-Driven Recognition, Reconstruction, Refinement, and Redeployment for Continual Robot Self-Improvement
- Nanyang Technological University(南洋理工大学)
- Xi’an Jiaotong University(西安交通大学)
- Dexmal
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
F4R提出失败驱动的真实-仿真-真实闭环框架,通过自动识别、重建失败场景并强化学习精炼策略,实现机器人持续自我改进,在OOD条件下超越基线18.75个百分点。
AI中文摘要:
当前视觉-语言-动作模型在现实世界中的性能从根本上受到专家演示覆盖范围有限及其对物理交互理解不足的制约。一种常见的补救措施是收集新遇到失败场景的额外现实世界演示。然而,这一过程成本高昂、效率低下、可能存在安全隐患且难以扩展。为应对这一挑战,我们提出了失败驱动崛起(F4R),一个失败驱动的真实-仿真-真实闭环学习框架,将现实世界失败转化为有针对性的策略改进。F4R首先使用智能体自动从轨迹回放中识别并诊断失败。它将每个失败重建为一个交互式、以物体为中心、保留任务相关空间与物理条件的桌面环境。随后,策略通过失败条件化的仿真-真实协同训练以及在重建环境中进行的目标强化学习得到精炼。改进后的策略随后被重新部署,同时新观察到的失败会持续反馈到下一轮重建与学习循环中。在四个操作任务上的现实世界评估表明,F4R实现了93.75%的分布内成功率和90.0%的分布外(OOD)成功率,在OOD条件下比预算匹配的目标行为克隆基线高出18.75个百分点,且无需收集额外的现实世界纠正演示。
英文摘要:
The real-world performance of current vision-language-action models is fundamentally constrained by the limited coverage of expert demonstrations and their insufficient understanding of physical interactions. A common remedy is to collect additional real-world demonstrations of newly encountered failures. However, this process is costly, inefficient, potentially unsafe, and difficult to scale. To address this challenge, we propose Failure for Rising (F4R), a failure-driven real-to-sim-to-real closed-loop learning framework that converts real-world failures into targeted policy improvement. F4R first uses an agent to automatically identify and diagnose failures from rollouts. It reconstructs each failure as an interactive, object-centric table-top environment that preserves the task-relevant spatial and physical conditions. The policy is then refined through failure-conditioned sim-real co-training followed by targeted reinforcement learning in the reconstructed environments. The improved policy is subsequently redeployed, while newly observed failures are continuously fed back into the next reconstruction and learning cycle. Real-world evaluations on four manipulation tasks show that F4R achieves 93.75% In-Distribution and 90.0% Out-of-Distribution (OOD) success, outperforming the budget-matched Targeted BC baseline by 18.75 percentage points under OOD conditions without collecting additional real-world corrective demonstrations.