发表机构
Shanghai AI Laboratory(上海人工智能实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
InternW0-$\Delta$提出统一世界动作模型,融合视觉动力学与动作生成,利用20K+小时异构开放数据预训练,在仿真和真实机器人上超越先前方法。
AI 中文摘要
世界动作模型(WAMs)联合建模视觉动力学与动作生成,用于通用机器人操作。一个核心挑战是将大规模预训练模型中的先验知识——包括视觉动力学、场景语义、几何和运动——整合到一个统一的框架中,以生成机器人动作。我们提出了InternW0-$\Delta$,一个在异构语料库上预训练的统一WAM,在仿真基准和真实机器人平台上均优于先前方法。InternW0-$\Delta$在Mixture-of-Transformers(MoT)框架内结合了预训练的视觉动力学、场景级语义、4D几何与运动先验以及动作生成。一个预训练的视频专家和一个动作专家在冻结的VLM的语义引导下交互,而一个预训练的4D基础模型通过仅训练时的蒸馏注入几何和运动先验。我们进一步引入了Causal Imprint,它从仅训练时的未来监督中学习与未来相关的场景变化,并在推理时无需未来视频展开即可直接向动作专家提供预测性表示。为了大规模联合训练,我们构建了一个包含机器人演示、UMI数据、第一人称人类演示和Ego2Robot数据的异构语料库,并在统一的状态-动作表示下进行整理和对齐。由此产生的语料库包含超过20K小时的处理后训练数据,据我们所知,这是同类中最大的开源语料库。我们在此语料库上预训练了InternW0-$\Delta$,并在仿真基准和真实机器人平台上展示了强大的性能。我们将开源训练代码、模型权重、基础设施、数据处理流程以及在许可允许范围内的处理后数据。项目页面:此https URL。
英文摘要
World Action Models (WAMs) jointly model visual dynamics and action generation for generalist robot manipulation. A central challenge is to integrate priors from large-scale pretrained models---including visual dynamics, scene semantics, geometry, and motion---into a unified framework for robot action generation. We introduce InternW0-$Δ$, a unified WAM pretrained on a heterogeneous corpus that outperforms prior methods across simulation benchmarks and real-robot platforms. InternW0-$Δ$ combines pretrained visual dynamics, scene-level semantics, 4D geometric and motion priors, and action generation within a Mixture-of-Transformers (MoT) framework. A pretrained video expert and an action expert interact under semantic guidance from a frozen VLM, while a pretrained 4D foundation model injects geometric and motion priors through training-only distillation. We further introduce Causal Imprint, which learns future-relevant scene changes from training-only future supervision and provides predictive representations directly to the action expert without future-video rollout at inference. For large-scale joint training, we construct a heterogeneous corpus of robot demonstrations, UMI data, egocentric human demonstrations, and Ego2Robot data, curated and aligned under a common state-action representation. The resulting corpus contains over 20K hours of processed training data, to our knowledge the largest open-source corpus of its kind. We pretrain InternW0-$Δ$ on this corpus and demonstrate strong performance across simulation benchmarks and real-robot platforms. We will open source the training code, model weights, infrastructure, data-processing pipeline, and processed data where licenses permit. Project page: https://internrobotics.github.io/InternW0-Delta/