发表机构
New York University Abu Dhabi; Chatsign Technology(纽约大学阿布扎比分校; Chatsign Technology)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出PEWAM模型,将机器人演示编译为效果程序,通过流匹配生成执行,实现可编辑、跨平台迁移,在多个基准上超越现有方法。
AI 中文摘要
机器人演示在相同的帧中记录了两件事:物体发生了什么变化,以及一个特定的机械臂如何实现这一变化。我们以第一项为条件。演示被编译成一个效果程序:移动物体的3D关键点轨迹、标记每个物体被握住位置的两个点、以及场景结束时的配置,同时移除演示者。PEWAM,一个7150万参数的世界-动作模型,生成效果、机器人执行、动作和终端状态作为四个流,具有独立的流匹配时间,因此固定一个程序并采样执行将推理转化为编程,从实时场景闭环重新求解。在保留的LIBERO-Goal任务上,一个演示的程序完成了90个回合中的40个,而相同的骨干网络在给定目标图像或语言以及已发布的演示条件方法下,最多完成19个;在Meta-World的三个保留类别上,它超过了最佳已发布结果,尽管在五类平均值上未超过。由于程序是一组坐标,人类可以编辑它:放置遵循移动的终端状态,抓取随旋转的接触点转动。同一程序在四个机器人臂上无需重新训练即可运行,并且在推动后,重新求解完成了60个回合中的23个,而重放演示完成了6个。在Franka臂上,针对其他任务的真实演示进行微调,从单个人类视频编译的程序完成了40次试验中的36次,而相同骨干网络以视频最后一帧作为目标图像的条件化完成了22次。
英文摘要
A robot demonstration records two things in the same frames: what happened to the objects, and how one particular arm made it happen. We condition on the first. A demonstration is compiled into an effect program: the 3D keypoint trajectories of the objects that moved, two points marking where each was held, and the configuration the scene ends in, with the demonstrator removed. PEWAM, a 71.5M-parameter world-action model, generates effect, robot execution, action and terminal state as four streams with independent flow-matching times, so clamping a program and sampling the execution turns inference into programming, re-solved closed loop from the live scene. On held-out LIBERO-Goal tasks, one demonstration's program completes 40 of 90 episodes, where the same backbone given a goal image or language, and published demonstration-conditioned methods, complete at most 19; on three of Meta-World's held-out classes it exceeds the best published results, though not on the five-class mean. Because a program is a set of coordinates, a person can edit it: the placement follows a shifted terminal state and the grasp turns with rotated contact points. The same program runs on four robot arms without retraining, and after a push, re-solving completes 23 of 60 episodes where replaying the demonstration completes 6. On a Franka arm, fine-tuned on real demonstrations of other tasks, programs compiled from single human videos complete 36 of 40 trials, against 22 for the same backbone conditioned on the video's last frame as a goal image.