发表机构
University of Science and Technology of China; Suzhou Artificial Intelligence Laboratory(中国科学技术大学; 苏州人工智能实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出UMR统一动作表示,分解为世界流与自我轨迹,通过WEPVLA策略实现从人类演示到异构机器人的零样本技能迁移,在仿真和真实实验中均取得高成功率。
AI 中文摘要
通用具身操作依赖于统一的动作表示,该表示能够跨具身泛化并易于扩展。然而,现有策略依赖于具身特定的动作空间,这使得跨具身演示难以大规模利用,并限制了向新具身和空间变体的迁移。为此,我们引入了通用操作表示(UMR),这是一种统一的动作表示,能够实现从人类演示到异构机器人的零样本技能迁移。UMR将操作分解为两个功能上不同但几何上关联的组件:与具身无关的世界流(World Flow),描述世界坐标系中与任务相关的物体运动;以及自我轨迹(Ego Trajectory),表示相对于当前姿态的末端执行器运动。我们将UMR实例化为世界-自我点VLA(WEPVLA),这是一个紧凑的0.5B参数策略,通过双流点动作适配器和统一的点动作专家在统一的几何动作空间中学习,并通过SE(3)共轭将两个组件耦合。为了提高数据效率,我们为UMR补充了数据高效策略(DES),该策略通过阶段感知的点云编辑来多样化物体配置,同时保留演示的接触几何。在仿真中,WEPVLA在LIBERO上平均成功率达到97.5%,在10任务RLBench基准上达到85.7%。在真实世界实验中,一个在人类演示上训练并由DES增强的策略能够零样本迁移到多样化的部署条件。每个任务收集约10分钟的人类演示且无需机器人演示,该策略在六个评估设置中平均成功率达到91.7%,而HumanEgo为60.8%。代码和附加材料可在该https URL获取。
英文摘要
General-purpose embodied manipulation hinges on a unified action representation that generalizes across embodiments and scales readily. Yet existing policies rely on embodiment-specific action spaces, making cross-embodiment demonstrations difficult to leverage at scale and limiting transfer to new embodiments and spatial variations. To this end, we introduce Universal Manipulation Representation (UMR), a unified action representation that enables zero-shot skill transfer from human demonstrations to heterogeneous robots. UMR decomposes manipulation into two functionally distinct yet geometrically linked components: embodiment-agnostic World Flow, which describes task-relevant object motion in the world frame, and Ego Trajectory, which represents end-effector motion relative to the current pose. We instantiate UMR as World--Ego Point VLA (WEPVLA), a compact 0.5B-parameter policy that learns in the unified geometric action space through a dual-stream Point Action Adapter and a unified Point Action Expert, with an $SE(3)$ conjugation coupling the two components. To improve data efficiency, we complement UMR with a Data-Efficient Strategy (DES) that diversifies object configurations through stage-aware point-cloud editing while preserving demonstrated contact geometry. In simulation, WEPVLA achieves average success rates of 97.5\% on LIBERO and 85.7\% on the 10-task RLBench benchmark. In real-world experiments, a single policy trained on human demonstrations augmented by DES transfers zero-shot to diverse deployment conditions. With about 10 minutes of collected human demonstrations per task and no robot demonstrations, it achieves 91.7\% average success across six evaluation settings, compared with 60.8\% for HumanEgo. Code and additional materials are available at https://umr-wepvla.github.io/.
CommentsSubmitted to IEEE International Conference on Robotics and Automation (ICRA)