发表机构
Carnegie Mellon University(卡内基梅隆大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出FAR框架,通过失败对比偏好适应和轻量动作扰动,使机器人从测试时失败中学习并自主恢复,结合成功轨迹持续改进策略,在仿真和真实任务中成功率分别提升17.6%和11.7%。
AI 中文摘要
机器人在真实环境中部署时不可避免地会遇到失败。简单的重试往往会重复相同的错误,而许多现有的恢复方法依赖于人工干预。在本文中,我们提出了失败感知重试(FAR)框架,该框架使机器人能够在测试时从先前的失败中学习,相应地调整其行为,并最终自主完成任务。FAR结合了失败对比偏好适应(从失败中构建偏好学习数据,引导策略远离先前不成功的行为)和重试过程中的轻量动作扰动(以鼓励局部探索)。我们进一步将成功的恢复轨迹纳入训练循环,以实现持续的策略改进。在仿真和真实世界操作任务中的实验表明,FAR显著提高了成功率和鲁棒性,在仿真中比标准扩散策略平均提升17.6%,在真实世界中提升11.7%。此外,在持续策略改进过程中,通过利用信息丰富的失败案例,FAR在重置预算和时间步预算下均显著提高了数据效率。
英文摘要
Robot policies inevitably encounter failures when deployed in real environments. Naive retries often repeat the same mistakes, while many existing recovery methods rely on human intervention. In this paper, we propose Failure-Aware Retry (FAR), a framework that enables robots to learn from previous failures at test time, adapt their behavior accordingly, and eventually complete the task autonomously. FAR combines Failure-Contrastive Preference Adaptation, which constructs preference learning data from failures to steer the policy away from previously unsuccessful behaviors, with lightweight action perturbations during retries to encourage local exploration. We further incorporate successful recovery trajectories into a training loop for continual policy improvement. Experiments in both simulation and real-world manipulation tasks show that FAR substantially improves success rates and robustness, with average gains of 17.6% over the standard diffusion policy in simulation and 11.7% in the real world. In addition, FAR improves data efficiency under both reset and timestep budgets during continual policy improvement by exploiting informative failure cases. Videos and code are available at https://hoar012.github.io/FAR-Project.
CommentsAccepted by CoRL 2026. Project Page: https://hoar012.github.io/FAR-Project