发表机构
Aether AI; University of California, San Diego; Vanderbilt University(Aether AI; 加利福尼亚大学圣迭戈分校; 范德堡大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出CausalWM,一个16B参数的具身世界模型,通过显式因果链式思维推理和31K小时数据的三阶段训练,在视频预测前捕捉物理因果依赖,并在多基准上达到最先进性能。
AI 中文摘要
具身世界模型从视觉观察和控制信号中学习预测未来的物理动态,其中物理知识隐式地纠缠在潜在表示中。我们提出了CausalWM,一个16B参数的具身世界模型,它在未来视频预测之前执行显式的因果链式思维推理。CausalWM将有用的变量组织成推理轨迹,使模型能够逐步捕捉物理演化背后的因果依赖关系。为了训练CausalWM,我们收集了31K小时的具身数据,并开发了一个三阶段范式,包括大规模视频预训练、因果CoT中期训练和多目标强化学习后期训练。尽管仅使用有限的监督CoT变量集,CausalWM展现出涌现的上下文学习能力,能够实现上下文视觉特征引导和高效的少步生成。CausalWM在语言条件、动作条件、单视图和多视图基准上均取得了最先进的性能,包括在TriWorldBench排行榜上排名第一。
英文摘要
Embodied world models learn to predict future physical dynamics from visual observations and control signals, where physical knowledge is implicitly entangled within latent representations. We introduce CausalWM, a 16B embodied world model that performs explicit causal chain-of-thought reasoning before future video prediction. CausalWM organizes useful variables into a reasoning trajectory, allowing the model to progressively capture causal dependencies underlying physical evolution. To train CausalWM, we collect 31K hours embodied data and develop a three-stage paradigm consisting of large-scale video pre-training, causal CoT mid-training, and multi-objective RL post-training. Despite using only a limited set of supervised CoT variables, CausalWM exhibits emergent in-context learning capabilities, enabling contextual visual feature guidance and efficient few-step generation. CausalWM achieves state-of-the-art performance across language-conditioned, action-conditioned, single-view and multi-view benchmarks, including Top-1 performance on TriWorldBench leaderboard.