发表机构
School of Information and Communication Engineering; Dalian University of Technology; Beta Infinity; Nanyang Technological University; Beijing Jiaotong University(信息与通信工程学院; 大连理工大学; 贝塔无限; 南洋理工大学; 北京交通大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对现有机器人操纵基准未评估故障恢复能力的问题,推出LIBERO-Recover基准,构建多级别故障场景并评估相关能力,推动具身智能体向故障恢复能力发展。
AI 中文摘要
视觉-语言-动作(VLA)或世界动作(WAM)模型近期在机器人操纵领域展现出卓越性能。在LIBERO基准上,SOTA方法已实现近100%的成功率,看似表明这些模型已准备好部署至现实世界。然而,现有基准上近乎完美的性能可能具有误导性:理想条件下的成功并不等同于现实世界的鲁棒性。现有基准主要评估预定义初始状态下的任务完成情况,而现实世界交互不可避免会出现抓取失败、碰撞、物体意外移动等故障。因此,机器人不仅要成功执行任务,还需识别故障并从中恢复以继续任务,但该能力目前基本未被衡量,这凸显了基准性能与现实世界可靠性之间的关键差距。为解决这一差距,我们推出LIBERO-Recover基准——一个面向机器人操纵故障恢复的大规模基准。该基准基于LIBERO构建,我们从SOTA具身模型中收集真实执行故障,构建了四个恢复级别共1000余个场景:(1)动作重试、(2)动作适配、(3)物体状态恢复、(4)环境恢复。我们评估四项核心能力:空间理解、物体结构推理、交互理解和拓扑推理。作为首个大规模具身故障恢复基准,LIBERO-Recover将评估从“机器人能否成功?”转向“机器人故障后能否恢复?”,以推动开发鲁棒且可泛化的具身智能体。该项目可在蓝色标注的此URL获取。
英文摘要
Vision-Language-Action (VLA) or World Action (WAM) models have recently demonstrated remarkable performance in robotic manipulation. On LIBERO, SOTA method have achieved nearly 100\% success rates, seemingly suggesting that the models are ready for deployment in real world. However, near perfect performance on existing benchmarks can be misleading: success under ideal conditions does not imply real world robustness. Existing benchmarks primarily evaluate task completion from predefined initial states, while real world interactions inevitably involve failures such as failed grasps, collisions, and unintended object movements. A robot must therefore not only execute tasks successfully, but also recognize and recover from failures to continue the task. Yet this capability remains largely unmeasured, revealing a critical gap between benchmark performance and real world reliability. To address this gap, we introduce LIBERO-Recover Benchmark, a large scale benchmark for failure recovery in robotic manipulation. Built upon LIBERO, we collect real execution failures from SOTA embodied models and construct 1,000+ scenarios across four recovery levels: (1) Action Retry, (2) Action Adaptation, (3) Object State Recovery, and (4) Environmental Recovery. We evaluate four core capabilities: spatial understanding, object structure reasoning, interaction understanding, and topological reasoning. As the first large-scale benchmark for embodied failure recovery, LIBERO-Recover shifts evaluation from \emph{Can the robot succeed?''} to \emph{Can the robot recover after failure?''}, promoting robust and generalizable embodied agents. The project will be avaible in \textcolor{blue}{https://liulin815.github.io/LIBERO-Recovery/}.