发表机构
Meta AI; University of Copenhagen; Physical Intelligence; Imperial College London(Meta AI; 哥本哈根大学; Physical Intelligence; 帝国理工学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出NAVA-WAM,通过从仅含观测的视频中直接预训练动作策略(原生动作先验学习),避免间接表征迁移,并在后训练中结合动作标签实现高效机器人控制,实验证明其优于现有方法且具备良好泛化性。
AI 中文摘要
世界动作模型将未来视觉动态与机器人动作预测相结合,但其可扩展性仍受限于对带有动作标注的机器人轨迹的需求。仅含观测的视频包含丰富的交互动态证据,但现有方法通常利用这些视频来预训练视觉表征(后续需适应控制任务),或推断潜在动作(随后将其映射到机器人指令)。我们提出NAVA-WAM,通过直接从仅含观测的视频中预训练动作策略,引入原生动作先验学习,避免了间接的表征到控制的迁移或独立的潜在动作模型。我们的训练包含两个阶段。首先,我们在仅含观测的视频上进行预训练,其中通过过渡结构的联合注意力将视觉转换上的未来视频流匹配监督传播,以优化Action-DiT并学习与动作相关的先验。其次,我们使用带有动作标签的演示数据,通过联合视频-动作流匹配对Action-DiT进行后训练以用于机器人控制,同时非对称注意力将视觉分支与迭代动作去噪解耦,实现高效的动作仅推理。大量实验表明,NAVA-WAM在分布内和分布外设置下均持续优于先前方法,同时展现出强大的动作标签效率和有效的真实机器人泛化能力。这些结果确立了原生动作先验学习作为一种有效方法,可直接从仅含观测的视频中预训练动作策略,为超越动作标注机器人数据提供了可扩展的路径。
英文摘要
World action models integrate future visual dynamics with robot action prediction, but their scalability remains limited by the need for action-annotated robot trajectories. Observation-only videos contain rich evidence about interaction dynamics, but existing approaches typically use them either to pretrain visual representations that must later be adapted for control, or to infer latent actions that are subsequently grounded to robot commands. We present NAVA-WAM, which introduces native action-prior learning by directly pretraining the action policy from observation-only videos, avoiding indirect representation-to-control transfer or a separate latent-action model. Our training consists of two stages. First, we pretrain on observation-only videos, where future-video flow-matching supervision over visual transitions is propagated through transition-structured joint attention to optimize the Action-DiT and learn action-relevant priors. Second, we use action-labeled demonstrations to post-train the Action-DiT for robot control through joint video--action flow matching, while asymmetric attention decouples the visual branch from iterative action denoising and enables efficient action-only inference. Extensive experiments show that NAVA-WAM consistently outperforms prior approaches under both in-distribution and out-of-distribution settings, while demonstrating strong action-label efficiency and effective real-robot generalization. These results establish native action-prior learning as an effective approach to directly pretrain action policies from observation-only videos, providing a scalable path beyond action-labeled robot data.
CommentsProject Page: https://zhaochongan.github.io/projects/NAVA-WAM