arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

从世界模型到世界动作模型:面向机器人学的简明教程

From World Models to World Action Models: A Concise Tutorial for Robotics

Xiaoxiong Zhang, Xiong Zeng, Wei Zhang

arXiv 2607.00836首次发表:更新:

AI 中文总结

本教程提出世界模型作为动作条件预测模型的设计空间视图,分类为观测空间和状态空间模型,并引入世界动作模型,总结四种代表性范式,为具身预测与控制提供结构化分类。

AI 中文摘要

世界模型越来越多地用于具身智能和生成式仿真,但其范围在不同社区中仍然模糊。本教程提出了世界模型作为动作条件预测模型的设计空间视图,这些模型估计任务相关观测或状态的未来演化。我们将现有方法分类为观测空间和状态空间世界模型,比较它们在视觉保真度、空间结构、物理可解释性和控制可用性方面的权衡。我们进一步引入了世界动作模型,它将预测的未来与可执行的机器人动作连接起来,并总结了四种代表性范式:想象-然后-执行、视频特征条件动作预测、联合视频-动作建模以及用于策略学习的辅助视频预测。本教程的目标是澄清世界(动作)模型的概念范围,并为具身预测和控制提供结构化的分类。

英文摘要

Rather than providing an exhaustive survey, this paper presents a concise tutorial on world models and world action models for robotics. After reading the tutorial, readers should have a clear understanding of what constitutes a "world", how world models and world action models are defined, and what roles they play within robotic AI systems. The tutorial also develops a unified perspective for comparing representative approaches, such as World Labs' spatial intelligence models, Yann LeCun's JEPA framework, and NVIDIA's Cosmos platform, and clarifies how these models differ in their representations, predictive capabilities, and interaction mechanisms.

CommentsGithub page: https://github.com/clearlab-sustech/WorldModelSurvey

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑