arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Magic-W0:面向物理智能的结构化世界-动作基础模型

Magic-W0: A Structured World-Action Foundation Model for Physical Intelligence

Xuhua Chen, Zhenhan Yin, Yuan Zhang, Lingfeng Zhang, He Zheng, Tong Mu, Shun Zuo, Dian Zhou, Di Wu, Xuan Zhou, Shaojie Wan, Rongtian Shen, Qiulong Xu, Yiduo Li, Yinglong Wang, Yanqian Wang, Kun Wang, Tao Zhang

arXiv 2609.39870首次发表:更新:

发表机构

Magiclab Robotics Inc(Magiclab机器人公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

Magic-W0通过结构化世界转移与层对齐交互架构,联合建模物理状态演化与连续动作,在RoboDojo-Sim取得最高得分27.10,并在真实机器人任务中展现强泛化与快速适应能力。

AI 中文摘要

世界-动作模型(WAMs)通过动作条件的环境动力学增强机器人策略,然而现有方法主要依赖于未来观测重建或通用潜在预测,缺乏与动作生成紧密耦合的结构化、面向控制的世界表示。我们提出Magic-W0,一种联合建模结构化物理状态演化与连续动作的世界-动作基础模型。Magic-W0将交互表示为结构化世界转移,包括当前状态、转移和未来状态。当前状态结合视觉-语言上下文与当前3D几何;转移由3D运动表示,捕获动作引起的三维变化;未来状态由未来语义表示,描述任务相关结果。为耦合预测与控制,我们提出层对齐的世界-动作交互架构,其中演化的动作假设条件化世界转移预测,而预测的世界表示持续为动作生成提供信息。Magic-W0在大规模以自我为中心的类人操作、UMI、真实机器人和仿真数据上预训练,利用预训练视觉模型对几何、3D运动和未来语义进行潜在监督。推理时干预表明,结构化世界表示对候选动作的变化作出系统性响应,且动作相关信息通过共享3D表示传播到未来语义预测中。在RoboDojo-Sim上,Magic-W0取得平均得分27.10,在比较的WAMs中最高。在多个真实机器人任务中,经过有限下游数据微调后也展现出强大的下游性能,支持泛化与快速适应。

英文摘要

World-action models (WAMs) augment robot policies with action-conditioned environment dynamics, yet existing approaches largely rely on future observation reconstruction or generic latent prediction and lack structured, control-oriented world representations tightly coupled with action generation. We introduce Magic-W0, a world-action foundation model that jointly models structured physical state evolution and continuous actions. Magic-W0 represents interaction as a Structured World Transition consisting of Current State, Transition, and Future State. Current State combines vision-language context with Current 3D Geometry; Transition is represented by 3D Motion capturing action-induced three-dimensional changes; and Future State is represented by Future Semantics describing task-relevant outcomes. To couple prediction and control, we propose a layer-aligned world-action interaction architecture in which evolving action hypotheses condition world-transition prediction, while predicted world representations continuously inform action generation. Magic-W0 is pre-trained on large-scale egocentric human manipulation, UMI, real-robot, and simulation data, with latent supervision for geometry, 3D motion, and future semantics from pre-trained visual models. Inference-time interventions show that structured world representations respond systematically to changes in candidate actions and that action-related information propagates through shared 3D representations into future semantic predictions. On RoboDojo-Sim, Magic-W0 achieves an average Score of 27.10, the highest among the compared WAMs. Across multiple real-robot tasks, it also demonstrates strong downstream performance after fine-tuning with limited downstream data, supporting generalization and rapid adaptation.

Comments29 pages, 15 figures, 7 tables. Project page: https://embodied.magiclab.top/works/wam/magic-w0/index.html; Code: https://github.com/MagiclabRobotics/Magic-W0

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑