LD4WAM:从人类视频学习世界动作模型的潜在动力学
LD4WAM: Learning Latent Dynamics from Human Videos for World Action Models
浏览论文内容
中文总结 AI 辅助
LD4WAM通过运动对齐的潜在动力学,结合语义重建与真实运动对齐训练的潜在动力学模型和MoT架构的世界动力学动作模型,在RoboTwin仿真及真实机器人上表现良好,可泛化到未见物体与背景。
中文摘要 AI 辅助
人类视频因相对于遥操作机器人数据具有多样性和采集成本低的优势,在训练世界动作模型(World Action Models,WAMs)中扮演着日益核心的角色。然而,大多数WAM仅通过预测像素级未来帧从这类视频中学习,得到的动力学无法直接执行;而运动重定向虽能恢复可直接执行的动作,但在不同 embodiment(实体形态)间存在较大视觉差距。因此,我们提出运动对齐的潜在动力学作为一种与实体形态无关的表示,以弥合视频先验与低级动作之间的鸿沟。我们进一步提出LD4WAM,它将通过语义重建和真实运动对齐训练的潜在动力学模型,与作为混合Transformer(mixture-of-transformers,MoT)构建的世界动力学动作模型配对,该模型保留完整的未来视频生成能力,并使用可学习查询从生成的未来中提取这些潜在动力学以用于动作条件设置。LD4WAM在我们整理的包含超过5000小时人类和机器人数据的统一数据集上进行预训练后,在RoboTwin仿真环境以及配备夹爪和灵巧手的真实机器人上表现出色,同时能很好地泛化到未见过的物体和背景。
英文摘要
Human video is playing an increasingly central role in training World Action Models (WAMs), owing to its diversity and low collection cost relative to teleoperated robot data. However, most WAMs learn from such video only by predicting pixel-level future frames, giving dynamics that are not directly actionable, whereas motion retargeting recovers directly actionable actions but leaves a large visual gap across embodiments. We therefore propose motion-aligned latent dynamics as an embodiment-agnostic representation to bridge video priors and low-level actions. We further present LD4WAM, which pairs a Latent Dynamics Model trained with semantic reconstruction and real motion alignment with a World Dynamics Action Model built as a mixture-of-transformers (MoT), which preserves full future-video generation and uses learnable queries to distill these latent dynamics from generated futures for action conditioning. Pretrained on our curated unified dataset of over 5{,}000 hours of human and robot data, LD4WAM performs strongly in RoboTwin simulation and on real robots equipped with both grippers and dexterous hands, while generalizing well to unseen objects and backgrounds.