发表机构
CMU; Columbia; NVIDIA(卡内基梅隆大学; 哥伦比亚大学; 英伟达)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
PointZero提出以三维点轨迹补全为预训练目标,无需机器人数据即可学习可迁移三维动力学,并在下游动作预测和模仿学习任务中优于基线。
AI 中文摘要
世界模型赋予感知系统预测场景在交互下如何演变的能力。当在多样化的数据量上进行训练时,它们最为有益,能为下游应用注入丰富的先验知识。现有方法通常需要机器人动作标签来学习动作条件下的三维动力学,这排除了网络视频数据进入训练池。我们研究将三维点轨迹补全作为预训练目标,以在无需机器人数据的情况下学习可迁移的三维动力学。给定单个RGB-D观测和稀疏的部分三维轨迹(轨迹),我们预测所有观测点的未来三维轨迹。我们表明,该目标能产生丰富的三维动力学先验,且无需机器人动作标签。我们贡献了一个包含290万合成帧的多样化数据集,涵盖可变形、铰接和刚体物体,并用其训练PointZero。我们证明,一个灵活且富有表现力的Transformer——PointZero,在同一数据上优于先前方法。我们通过将PointZero用于两个下游应用的后期训练来展示我们预训练目标的实用性:(1)动作条件下的三维动力学预测和(2)模仿学习。当微调以末端执行器姿态为条件时,PointZero在最近的PGND三维动力学基准上优于基线。当微调以预测机器人动作和三维轨迹时,PointZero在6/7个模拟和真实世界机器人操作任务上优于或匹配基线。我们进一步评估从头训练PointZero,以将我们提出的架构的益处与我们提出的预训练目标和数据集的益处分离开来。我们发布了数据集、检查点和完整的训练方案。
英文摘要
World models endow perceptual systems with the ability to predict how scenes evolve under interaction. They are most beneficial when trained on diverse volumes of data, to instill a rich prior into downstream applications. Existing methods typically require robot action labels to learn action-conditioned 3D dynamics, which excludes web video data from the training pool. We study 3D point track completion as a pre-training objective for learning transferable 3D dynamics without robot data. Given a single RGB-D observation and sparse partial 3D trajectories (tracks), we predict future 3D tracks of all observed points. We show this objective produces a rich 3D dynamics prior, without requiring robot action labels. We contribute a diverse dataset of 2.9 million synthetic frames spanning deformable, articulated, and rigid objects, and use it to train PointZero. We show that a flexible and expressive transformer, PointZero, outperforms prior methods on the same data. We demonstrate the utility of our pre-training objective by post-training PointZero for two downstream applications: (1) action-conditioned 3D dynamics prediction and (2) imitation learning. When fine-tuned to condition on end-effector pose, PointZero outperforms the baselines on the recent PGND 3D dynamics benchmark. When fine-tuned to predict robot actions and 3D tracks, PointZero outperforms or matches the baselines on 6/7 simulated and real-world robot manipulation tasks. We furthermore evaluate training PointZero from scratch to isolate the benefits of our proposed architecture from those of our proposed pre-training objective and dataset. We release the dataset, checkpoints, and full training recipe.
Commentshttps://pointzero-wm.github.io/