发表机构
University of Minnesota(明尼苏达大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
PointCast提出一个点集世界模型,用扩散变换器预测物体点轨迹,统一处理刚体、铰接体和可变形物体操作,在模拟和真实数据上均优于基线。
AI 中文摘要
世界模型对机器人操作非常有用,因为机器人可以在执行动作之前预测动作如何改变物体的状态。我们提出了PointCast,一个点集世界模型,涵盖刚体、铰接体和可变形物体的操作。其状态是物体和末端执行器上的3D点集,无需网格且与拓扑无关。每个点保持其身份,并在其自身轨迹上进行监督,这教会模型每个点的去向,而不仅仅是点所形成的形状。其主干是一个扩散变换器,对短窗口内的未来点位置进行去噪,以点的近期历史和指令的末端执行器运动为条件。主干的注意力在局部和全局之间交替,对末端执行器的交叉注意力承载耦合。这一架构(19.8M参数)和单一训练方案覆盖四个领域:刚体、布料、绳索和多关节柜,每个领域训练一个单独的检查点。在随机化模拟上训练,并在相同指标上对四个基线进行评分,它在四个领域中三个表现最佳,在刚体上排名第二。在真实机器人遥操作数据集上训练,它在六个类别中四个具有最低平均误差,在其他两个中排名第二,并在所有六个类别中优于数据集自身的模型;零样本情况下,其模拟检查点在四个捕获中两个最佳。在基于采样的模型预测控制中冻结,每个窗口一次网络评估,它在64个回合中规划四个模拟任务,在每个任务上与所有基线竞争或优于它们。项目网站:https://pointcast-wm.github.io。
英文摘要
World models are useful for robotic manipulation because robots can predict how actions change the states of objects before executing them. We present PointCast, a point-set world model that spans rigid, articulated, and deformable object manipulation. Its state is a set of 3D points on the object and the end-effector, mesh-free and topology-agnostic. Each point keeps its identity and is supervised on its own trajectory, which teaches the model where every point goes rather than only the shape the points form. Its backbone is a diffusion transformer that denoises a short window of future point positions, conditioned on the points' recent history and the commanded end-effector motion. The backbone's attention alternates between local and global, and cross-attention to the end-effector carries the coupling. This one architecture at 19.8M parameters and one training recipe cover four regimes, rigid objects, cloth, rope, and multi-joint cabinets, with a separate checkpoint trained for each. Trained on randomized simulation and scored against four baselines on the same metric, it is best on three of four regimes and second on rigid. Trained on a real-world robot teleoperation dataset, it has the lowest mean error in four of its six categories, is second in the other two, and improves on the dataset's own model in all six; zero-shot, its simulation checkpoints are best on two of four captures. Frozen inside sampling-based model-predictive control at one network evaluation per window, it plans four simulated tasks over 64 episodes, competitive with or outperforming every baseline on each. Project website at https://pointcast-wm.github.io.
Comments8 pages, 8 figures, 5 tables. Project page: https://pointcast-wm.github.io. This work has been submitted to the IEEE for possible publication. Copyright may be transferred without notice, after which this version may no longer be accessible