发表机构
Magiclab Robotics Inc(Magiclab机器人公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
Magic-W0通过结构化世界转移与层对齐交互架构,联合建模物理状态演化与连续动作,在RoboDojo-Sim取得最高得分27.10,并在真实机器人任务中展现强泛化与快速适应能力。
AI 中文摘要
世界-动作模型(WAMs)通过动作条件的环境动力学增强机器人策略,然而现有方法主要依赖于未来观测重建或通用潜在预测,缺乏与动作生成紧密耦合的结构化、面向控制的世界表示。我们提出Magic-W0,一种联合建模结构化物理状态演化与连续动作的世界-动作基础模型。Magic-W0将交互表示为结构化世界转移,包括当前状态、转移和未来状态。当前状态结合视觉-语言上下文与当前3D几何;转移由3D运动表示,捕获动作引起的三维变化;未来状态由未来语义表示,描述任务相关结果。为耦合预测与控制,我们提出层对齐的世界-动作交互架构,其中演化的动作假设条件化世界转移预测,而预测的世界表示持续为动作生成提供信息。Magic-W0在大规模以自我为中心的类人操作、UMI、真实机器人和仿真数据上预训练,利用预训练视觉模型对几何、3D运动和未来语义进行潜在监督。推理时干预表明,结构化世界表示对候选动作的变化作出系统性响应,且动作相关信息通过共享3D表示传播到未来语义预测中。在RoboDojo-Sim上,Magic-W0取得平均得分27.10,在比较的WAMs中最高。在多个真实机器人任务中,经过有限下游数据微调后也展现出强大的下游性能,支持泛化与快速适应。
英文摘要
World-action models (WAMs) augment robot policies with action-conditioned environment dynamics, yet existing approaches largely rely on future observation reconstruction or generic latent prediction and lack structured, control-oriented world representations tightly coupled with action generation. We introduce Magic-W0, a world-action foundation model that jointly models structured physical state evolution and continuous actions. Magic-W0 represents interaction as a Structured World Transition consisting of Current State, Transition, and Future State. Current State combines vision-language context with Current 3D Geometry; Transition is represented by 3D Motion capturing action-induced three-dimensional changes; and Future State is represented by Future Semantics describing task-relevant outcomes. To couple prediction and control, we propose a layer-aligned world-action interaction architecture in which evolving action hypotheses condition world-transition prediction, while predicted world representations continuously inform action generation. Magic-W0 is pre-trained on large-scale egocentric human manipulation, UMI, real-robot, and simulation data, with latent supervision for geometry, 3D motion, and future semantics from pre-trained visual models. Inference-time interventions show that structured world representations respond systematically to changes in candidate actions and that action-related information propagates through shared 3D representations into future semantic predictions. On RoboDojo-Sim, Magic-W0 achieves an average Score of 27.10, the highest among the compared WAMs. Across multiple real-robot tasks, it also demonstrates strong downstream performance after fine-tuning with limited downstream data, supporting generalization and rapid adaptation.
Comments29 pages, 15 figures, 7 tables. Project page: https://embodied.magiclab.top/works/wam/magic-w0/index.html; Code: https://github.com/MagiclabRobotics/Magic-W0