发表机构
University of California, San Diego; University of Texas at Austin; NVIDIA(加州大学圣迭戈分校; 德克萨斯大学奥斯汀分校; 英伟达)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
Recova通过数字孪生中智能体引导的故障恢复,结合DAgger训练,将操作失败转化为可重用技能,在LIBERO-Pro和MolmoSpaces上分别达到78.8%和64.9%的成功率,并大幅减少人工干预。
AI 中文摘要
操作失败可能使场景处于任务策略无法恢复的状态。学习纠正行为需要可扩展的故障探索和物理基础。我们提出Recova,一个智能体引导框架,在重建的数字孪生中联合开发任务执行和恢复,然后通过真实世界经验验证并优化两者。在孪生中,智能体诊断故障、测试纠正程序,并为单独的策略收集成功的任务和恢复轨迹。在部署期间,它监控进度,调用学习或程序化恢复,验证场景恢复,并恢复执行。当没有合适的恢复可用时,人类演示解决故障并进入学习循环,使系统能够扩展其恢复能力。物理轨迹和人类演示被路由到相应的策略进行DAgger训练。在六个LIBERO-Pro设置和四个MolmoSpaces类别中,Recova实现了78.8%和64.9%的平均成功率,而最强基线的成功率分别为71.7%和38.0%。通过在四个真实机器人工作站上并行收集,DAgger微调将平均成功率从23.8%提高到77.5%,恢复技能进一步将其提高到87.5%。在一个任务的四轮收集中,观察到的人工干预从87.5%下降到0%。这些结果共同表明,智能体引导的恢复如何将失败转化为可重用的能力,提高鲁棒性,同时逐步减少人工干预。项目页面:此https URL
英文摘要
Manipulation failures can leave scenes in states from which a task policy cannot recover. Learning corrective behaviors requires scalable failure exploration and physical grounding. We present Recova, an agent-guided framework that jointly develops task execution and recovery in a reconstructed digital twin, then verifies and refines both through real-world experience. In the twin, the agent diagnoses failures, tests corrective programs, and collects successful task and recovery rollouts for separate policies. During deployment, it monitors progress, invokes a learned or programmatic recovery, verifies scene restoration, and resumes execution. When no suitable recovery is available, a human demonstration resolves the failure and enters the learning loop, allowing the system to expand its recovery capabilities. Physical rollouts and human demonstrations are routed to the corresponding policy for DAgger training. Across six LIBERO-Pro settings and four MolmoSpaces categories, Recova achieves 78.8% and 64.9% mean success, compared with 71.7% and 38.0% for the strongest baselines. With parallel collection across four real-robot workstations, DAgger fine-tuning raises mean success from 23.8% to 77.5%, and recovery skills further raise it to 87.5%. Over four collection rounds on one task, observed human intervention falls from 87.5% to 0%. Together, these results show how agent-guided recovery turns failures into reusable capabilities, improving robustness while progressively reducing human intervention. Project page: https://www.liuisabella.com/Recova
CommentsProject page: https://www.liuisabella.com/Recova