arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

从世界模型到世界行动模型:重新思考下一状态预测

From World Models to World Action Models: Rethinking Next-State Prediction

Tingyu Yuan, Ziming Ji, Biaoliang Guan, Wen Ye, Wenrui Tian, Zhaopeng Gu, Feihong Zhang, Xu Yang, Yan Huang, Zhaowen Li, Chaoyang Zhao, Jinqiao Wang

arXiv 2609.34414首次发表:更新:

发表机构

CASIA; UCAS; BUPT; XJTU; WHU; THU; Yinwang Intelligent Technology Co. Ltd.(中国科学院自动化研究所; 中国科学院大学; 北京邮电大学; 西安交通大学; 武汉大学; 清华大学; 银旺智能科技有限公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出CF-WAM动态下一状态预测框架,通过多投影统一监督改进世界行动模型,提升训练效率与跨具身泛化,在多个基准上达到最优性能。

AI 中文摘要

预测下一状态是世界模型建模物理动力学的核心范式,强调预测保真度。随着世界模型演变为世界行动模型(WAMs),现有方法仍将训练前的下一状态固定为RGB、单一潜在特征或预定义目标的静态组合,从而将行动学习限制在特定表示所保留的归纳偏差内。为解决这一局限,我们提出CF-WAM,一种动态下一状态预测框架,该框架对同一未来的视觉、语义、几何和交互投影进行采样,将其标准化为统一的视频形式,并监督一个统一的WAM跨这些投影进行学习。这些投影暴露的行动相关约束在训练步骤中累积,迫使WAM捕获支持同一行动条件未来的多重投影的底层状态转移结构。这种动态机制还为人类和机器人学习提供了自然的跨具身动力学参考框架。通过跨不同下一状态参数化的联合学习,异构的人类和机器人经验可以绕过外观差异,直接贡献于共享的状态转移学习,提升跨具身泛化能力。实验表明,CF-WAM提高了训练效率和最终控制性能,同时有效将人类经验转化为策略增益。CF-WAM在RoboCasa-GR1上实现了最先进的性能,平均成功率为82.50%,在LIBERO-Plus上达到82.65%,在真实世界评估中高达84.00%。

英文摘要

Predicting the next state is a core paradigm of World Models for modeling physical dynamics, emphasizing prediction fidelity. As World Models evolve into World-Action Models (WAMs), existing methods still fix the next state before training as RGB, a single latent feature, or a static combination of predefined targets, thereby constraining action learning to the inductive biases preserved by a particular representation. To address this limitation, we propose CF-WAM, a dynamic next-state prediction framework that samples visual, semantic, geometric, and interaction projections of the same future, standardizes them into a common video form, and supervises a unified WAM across these projections. The action-relevant constraints exposed by these projections accumulate across training steps, forcing WAM to capture the underlying state-transition structure that supports multiple projections of the same action-conditioned future. This dynamic mechanism also provides a natural cross-embodiment dynamics reference frame for Human and Robot learning. By jointly learning across different next-state parameterizations, heterogeneous Human and Robot experience can bypass appearance differences and directly contribute to shared state-transition learning, improving cross-embodiment generalization. Experiments show that CF-WAM improves both training efficiency and final control performance, while translating Human experience effectively into policy gains. CF-WAM achieves state-of-the-art performance on RoboCasa-GR1 with an average success rate of 82.50%, while reaching 82.65% on LIBERO-Plus and up to 84.00% in real-world evaluations.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑