ManiUnit:面向长 horizon 任务的操作技能数据集与基准
ManiUnit: A Manipulation Skill Dataset and Benchmark for Long-Horizon Tasks
浏览论文内容
中文总结 AI 辅助
本研究针对长 horizon 移动操作技能学习与评估的挑战,构建了 ManiUnit 数据集与基准,实验表明基于该数据集训练的技能策略可显著提升操作成功率与完整任务成功率。
中文摘要 AI 辅助
长 horizon 移动操作要求机器人在多房间环境中导航,并根据单一自然语言指令执行一系列操作技能。学习和评估这些技能面临三大挑战:固定任务指令下的相似观测可能导致技能选择模糊;即使前序技能成功,下一技能继承的机器人状态也可能偏离其演示的起始状态,进而影响执行;任务级指标不利于技能层面的诊断,且早期失败会导致后续技能无法被测试。为此,我们引入 ManiUnit,一个基于 50 项 BEHAVIOR-1K 活动构建的操作技能数据集与基准。其数据集包含 21 类技能、417 个子任务下的 137899 个片段,基准包含 1260 个测试实例。对应地,ManiUnit 为每个片段配对明确的子任务指令;测量对机器人基座起始位置或关节配置扰动的敏感性;并恢复中间模拟器状态、定义局部成功条件,使每个技能无需执行前序阶段即可评估。对代表性视觉-语言-动作(VLA)策略的评估显示,相似的 aggregate 分数可能掩盖显著的技能间差异。测试的起始状态扰动也会降低执行效果:在完整基准上,关节扰动使成功率较演示起始状态的成功率降低约 56%。在两项长 horizon 活动中,基于 ManiUnit 片段训练的技能策略实现了 78.7% 的局部操作成功率,而基于完整演示训练的任务策略仅为 49.3%。这些训练后的技能进一步支持这些活动的完整任务执行,通过规划器协调任务与技能策略,将完整任务成功率从 4.0% 提升至 18.0%。
英文摘要
Long-horizon mobile manipulation requires a robot to navigate multi-room environments and execute a sequence of manipulation skills under a single natural language instruction. Learning and evaluating these skills present three challenges: similar observations under a fixed task instruction may make skill selection ambiguous; even when a preceding skill succeeds, the robot state inherited by the next skill may deviate from its demonstrated starting states and affect execution; and task-level metrics hinder skill-specific diagnosis, while early failures leave later skills untested. We therefore introduce ManiUnit, a manipulation skill dataset and benchmark built from 50 BEHAVIOR-1K activities. Its dataset contains 137,899 segments across 21 skill types and 417 subtasks, and its benchmark contains 1,260 test instances. Correspondingly, ManiUnit pairs each segment with an explicit subtask instruction; measures sensitivity to perturbations of the robot's starting base position or joint configuration; and restores intermediate simulator states and defines local success conditions so that each skill can be evaluated without executing preceding stages. Evaluations of representative vision-language-action (VLA) policies show that similar aggregate scores can hide substantial per-skill differences. The tested starting-state perturbations also degrade execution: on the full benchmark, joint perturbations reduce success rates by approximately 56% relative to those from demonstrated starting states. On two long-horizon activities, a skill policy trained on ManiUnit segments achieves 78.7% local manipulation success, compared with 49.3% for a task policy trained on complete demonstrations. The trained skills further support complete-task execution on these activities, as coordinating the task and skill policies through a planner raises full-task success from 4.0% to 18.0%.
发表机构
- Nanjing University of Science and Technology(南京理工大学)
- Northwestern Polytechnical University(西北工业大学)
- Intellifusion(云天励飞)
机构由 AI 辅助整理,请以论文原文为准。