arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

CausalWM:具身世界模型的因果链式思维推理

CausalWM: Causal Chain-of-Thought Reasoning for Embodied World Model

Ziming Xu, Shuang Liang, Ruobing Han, Ziqiao Xi, Mingxing Rao, Kun Zhou, Zijun Zhang, Yuchen Yan, Yufan Wei, Junbo Huang, Yifei Shao, Fang Nan, Biwei Huang

arXiv 2609.23184首次发表:更新:

发表机构

Aether AI; University of California, San Diego; Vanderbilt University(Aether AI; 加利福尼亚大学圣迭戈分校; 范德堡大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出CausalWM,一个16B参数的具身世界模型,通过显式因果链式思维推理和31K小时数据的三阶段训练,在视频预测前捕捉物理因果依赖,并在多基准上达到最先进性能。

AI 中文摘要

具身世界模型从视觉观察和控制信号中学习预测未来的物理动态,其中物理知识隐式地纠缠在潜在表示中。我们提出了CausalWM,一个16B参数的具身世界模型,它在未来视频预测之前执行显式的因果链式思维推理。CausalWM将有用的变量组织成推理轨迹,使模型能够逐步捕捉物理演化背后的因果依赖关系。为了训练CausalWM,我们收集了31K小时的具身数据,并开发了一个三阶段范式,包括大规模视频预训练、因果CoT中期训练和多目标强化学习后期训练。尽管仅使用有限的监督CoT变量集,CausalWM展现出涌现的上下文学习能力,能够实现上下文视觉特征引导和高效的少步生成。CausalWM在语言条件、动作条件、单视图和多视图基准上均取得了最先进的性能,包括在TriWorldBench排行榜上排名第一。

英文摘要

Embodied world models learn to predict future physical dynamics from visual observations and control signals, where physical knowledge is implicitly entangled within latent representations. We introduce CausalWM, a 16B embodied world model that performs explicit causal chain-of-thought reasoning before future video prediction. CausalWM organizes useful variables into a reasoning trajectory, allowing the model to progressively capture causal dependencies underlying physical evolution. To train CausalWM, we collect 31K hours embodied data and develop a three-stage paradigm consisting of large-scale video pre-training, causal CoT mid-training, and multi-objective RL post-training. Despite using only a limited set of supervised CoT variables, CausalWM exhibits emergent in-context learning capabilities, enabling contextual visual feature guidance and efficient few-step generation. CausalWM achieves state-of-the-art performance across language-conditioned, action-conditioned, single-view and multi-view benchmarks, including Top-1 performance on TriWorldBench leaderboard.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑