AI 中文总结
UniWAM 提出统一世界-动作模型,集成物理推理、世界生成与动作预测,通过自然语言动作表示和互补预训练,实现 SOTA 性能并揭示人机协同训练缩放定律。
AI 中文摘要
视觉-语言-动作模型受益于预训练视觉-语言模型的理解和推理能力,但仅依靠动作监督对世界动态的 grounding 有限。相反,世界-动作模型继承了视频生成模型的时空先验,但在分布偏移下的语义理解和推理方面仍受限。我们提出了 UniWAM,一种统一架构,集成了物理推理器、世界生成器和动作预测器,以联合学习物理世界的语义理解、视觉生成和动作预测。为确保训练数据质量,我们为人类自我中心数据和机器人数据开发了严格的数据清洗和标注流程。为使视觉-语言组件适应具身任务同时保留其继承的语言能力,我们以自然语言表示低级动作,并引入一种预训练方案,将来自视觉问答(VQA)数据、人类自我中心数据和机器人演示的互补监督分配给适当的模型组件。在后训练阶段,未来视觉噪声增强减少了对精确未来预测的依赖,而历史条件流匹配利用编码的动作历史来初始化动作生成。这些设计共同显著减少了去噪步骤,同时保持了性能。UniWAM 在多项评估中实现了最先进(SOTA)性能,包括分布内性能、鲁棒性、泛化、指令遵循和长时程任务执行。此外,我们发现了统一人机协同训练的对数线性缩放定律,证明了在大规模混合人类和机器人数据上进行预训练的有效性。
英文摘要
Vision-language-action models benefit from the understanding and reasoning capabilities of pretrained vision-language models, but action-only supervision provides limited grounding in world dynamics. Conversely, world-action models inherit spatiotemporal priors from video generation models, yet remain limited in semantic understanding and reasoning under distribution shifts. We introduce UniWAM, a unified architecture that integrates a physical reasoner, a world generator, and an action predictor to jointly learn semantic understanding of the physical world, visual generation, and action prediction. To ensure the quality of the training data, we developed a rigorous data cleaning and annotation pipeline for both human egocentric data and robot data. To adapt the vision-language component to embodied tasks while preserving its inherited language capabilities, we represent low-level actions in natural language and introduce a pre-training recipe that assigns complementary supervision from visual question answering (VQA) data, human egocentric data, and robot demonstrations to the appropriate model components. During post-training, future visual noise augmentation reduces reliance on precise future predictions, while history-conditioned flow matching uses encoded action history to initialize action generation. Together, these designs significantly reduce denoising steps while maintaining performance. UniWAM achieves state-of-the-art (SOTA) performance across multiple evaluations, including in-distribution performance, robustness, generalization, instruction following, and long-horizon task execution. Furthermore, we uncover a log-linear scaling law of unified human-robot co-training, demonstrating the effectiveness of large-scale pre-training on a mixture of human and robot data.