发表机构
Carnegie Mellon University(卡内基梅隆大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
ManiSkillFormer通过任务条件几何契约连接感知与行动,利用LLM生成技能模板,实现无演示的组合操作,在多种任务上取得高成功率。
AI 中文摘要
我们提出了ManiSkillFormer,一个用于无演示和组合式机器人操作的神经常规框架。ManiSkillFormer不学习端到端的视觉运动策略,而是引入任务条件几何契约,明确构建感知与行动之间的接口。每个操作技能声明执行所需的语义几何基元,如物体关键点和表面法线。基于人类定义的技能结构,LLM智能体为不同物体和任务上下文生成这些契约及相应的运动模板。这些契约引导感知模块从观测中提取任务相关的3D基元,进而用于实例化存储在技能库中的可复用运动模板。我们在Galaxea R1-Lite双臂机器人上评估了ManiSkillFormer,涵盖三种设置:8个物体类别、30个不同实例的零样本抓取放置任务;包括拧开、倒水、按压和折叠在内的功能性操作任务;以及3个长时程任务。ManiSkillFormer在平均成功率上高于评估的基线和两个消融管线:无演示抓取放置任务成功率为88.24%,功能性操作平均成功率为75.00%,长时程任务完成率为50%至80%。这些结果表明,我们的设计实现了跨物体和任务的组合式、可复用操作,无需针对每个物体进行策略微调或额外的机器人演示。
英文摘要
Adapting robotic manipulation to new objects and tasks often requires additional demonstrations or manual engineering. Reusable manipulation skills can reduce this effort, but adapting these skills to new scenes remains challenging. We present ManiSkillFormer, a framework for demonstration-free and compositional manipulation that connects perception and action through explicit geometric contracts. Building on reusable skill schemas, LLM agents generate contracts specifying the geometry primitives required by each skill, together with corresponding motion templates for semantic objects and task contexts. These contracts guide a perception module to ground task-relevant 3D geometry from observations, which is then used to instantiate reusable motion templates in a skill library. We evaluate ManiSkillFormer on a dual-arm robot across demonstration-free pick-and-place with 30 instances from 8 object categories, functional manipulation including unscrewing, pouring, pressing, and folding, and three long-horizon tasks. ManiSkillFormer achieves an average success rate of 88.97% for pick-and-place, 75.00% for functional manipulation, and completion rates of 50--80% across the long-horizon tasks, outperforming the evaluated baselines and two ablated pipelines. These results demonstrate the potential of explicit geometric contracts to support skill reuse and composition across objects and tasks without per-object policy fine-tuning or additional robot demonstrations. The project website is at https://patricia1019.github.io/ManiSkillFormer/.