发表机构
University of Calgary; Mila(卡尔加里大学; 米拉研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
REBOOT是首个以失败为核心信号的机器人操作基准,包含18个精密装配任务的2,160个演示,通过阶段级标注和恢复轨迹,揭示并评估策略在装配各阶段的失败模式。
AI 中文摘要
机器人学习策略会以特有的方式失败:在不确定状态下停滞,在接触密集的对齐过程中漂移,在精密任务中偏离目标数毫米。然而,训练数据集主要由成功演示组成,而现实世界的基准常常将性能简化为二元成功。这限制了对恢复的监督以及对失败发生位置的分析。我们提出REBOOT(Recovery Episode Benchmark for Off-nominal Trajectories,非标称轨迹恢复片段基准),这是第一个将失败作为首要信号的机器人操作基准。REBOOT包含18个精密装配任务中的2,160个演示,每个任务被分解为五个共享阶段:Align(pick)、Engage(pick)、Transport、Align(place)和Engage(place),支持超越终端成功的阶段级评估。失败在各阶段中被引入,并配以专家恢复轨迹,使系统返回有效的继续状态。任务通过匹配的安装-移除对标注了旋转对称性、配合间隙精度等级和装配方向。失败片段按阶段和分类失败模式进行标注,从而能够归因于运动学阶段和公差违反。数据包括来自四个视角的同步RGB-D观测,以及阶段级成功和失败条件的基于自然语言的描述。数据集的一半包含专家演示;另一半包含恢复演示,这些演示的采样反映了在模仿学习策略回滚中观察到的失败。我们使用阶段级完成率对动作分块变换器、扩散和π0-FAST策略进行基准测试,揭示了被二元评估隐藏的模型特定失败点。数据集和代码:此https URL
英文摘要
Robot learning policies fail in characteristic ways: they stall in uncertain states, drift during contact-rich alignment, and miss targets by millimetres in precision tasks. Yet training datasets consist largely of successful demonstrations, while real-world benchmarks often reduce performance to binary success. This limits both supervision for recovery and analysis of where failures occur. We introduce REBOOT (Recovery Episode Benchmark for Off-nominal Trajectories), the first robot manipulation benchmark designed around failure as a first-class signal. REBOOT contains 2,160 demonstrations across 18 precision assembly tasks, each decomposed into five shared phases: Align(pick), Engage(pick), Transport, Align(place), and Engage(place), enabling phase-level evaluation beyond terminal success. Failures are introduced across phases and paired with expert recovery trajectories that return the system to a valid continuation state. Tasks are annotated with rotational symmetry, engagement-clearance precision tier, and assembly direction through matched install-remove pairs. Failure episodes are labeled by phase and categorical failure mode, enabling attribution to kinematic stage and tolerance violation. Data includes synchronized RGB-D observations from four viewpoints and grounded natural-language descriptions of phase-level success and failure conditions. Half the dataset contains expert demonstrations; the other half contains recovery demonstrations sampled to reflect failures observed in imitation-learned policy rollouts. We benchmark action-chunked transformer, diffusion, and $π_0$-FAST policies using phase-level completion rates, revealing model-specific failure points hidden by binary evaluation. Dataset and code: https://nanayawoa.github.io/REBOOT
CommentsProject website: https://nanayawoa.github.io/REBOOT