arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

世界-动作模型用于机器人学习与控制:综述

World-Action Models for Robot Learning and Control: A Survey

Zuxing Lu, Hongjia Zhai, Guanzhi Wang, Huajian Zeng, Jiaqi Yang, Jingyu Liu, Lei Cheng, Yuantai Zhang, Yuheng Qiu, Zezhou Cheng, Ivan Laptev, Danfei Xu, Benjamin Riviere, Giuseppe Loianno, Eric Xing, Xingxing Zuo

arXiv 2609.16074首次发表:更新:

发表机构

MBZUAI; Caltech; Amazon FAR; University of Virginia; Georgia Tech; New York University; UC Berkeley(穆罕默德·本·扎耶德人工智能大学; 加州理工学院; 亚马逊FAR; 弗吉尼亚大学; 佐治亚理工学院; 纽约大学; 加州大学伯克利分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本综述系统梳理了世界-动作模型(WAMs),该模型结合未来世界预测与动作生成,通过统一分类法分析其方法、应用及挑战,旨在为具身机器人智能提供技术基础。

AI 中文摘要

在开放环境中运行的机器人在部分可观测性、物理约束和动态任务情境下行动。除了将观测和语言指令映射到动作之外,它们还必须预测候选动作可能如何影响未来状态和与任务相关的结果。世界模型、视频生成和视觉-语言-动作(VLA)策略的最新进展推动了世界-动作模型(WAMs)的发展,该模型将未来世界预测与可执行动作生成相结合。本综述提供了面向机器人的WAMs回顾。我们阐明了其相对于传统世界模型、基于模型的强化学习、动作条件视频生成和反应式VLA策略的范围,并通过一个统一的分类法组织现有方法,涵盖表示、转换建模、动作接口、架构、训练流程、数据模态和扩展策略。我们进一步回顾了WAMs在操作、导航和自动驾驶中的应用,并总结了用于评估WAM系统的数据集、基准、指标和协议。最后,我们讨论了动作对齐、世界-动作分解、空间和多视图一致性、长时记忆、用于闭环策略学习的神经模拟以及高效推理等关键挑战。总体而言,本综述旨在为将预测性世界建模与动作生成相结合提供简洁的技术基础,以实现更可靠的具身机器人智能。项目页面:此https URL。

英文摘要

Robots operating in open environments act under partial observability, physical constraints, and dynamic task contexts. Beyond mapping observations and language instructions to actions, they must anticipate how candidate actions may affect future states and task-relevant outcomes. Recent advances in world models, video generation, and Vision-Language-Action (VLA) policies have motivated the development of World-Action Models (WAMs), which couple future world prediction with executable action generation. This survey provides a robotics-oriented review of WAMs. We clarify their scope relative to conventional world models, model-based reinforcement learning, action-conditioned video generation, and reactive VLA policies, and organize existing methods through a unified taxonomy covering representations, transition modeling, action interfaces, architectures, training pipelines, data modalities, and scaling strategies. We further review applications of WAMs in manipulation, navigation, and autonomous driving, and we summarize the datasets, benchmarks, metrics, and protocols used to evaluate WAM systems. Finally, we discuss key challenges in action alignment, world-action factorization, spatial and multi-view consistency, long-horizon memory, neural simulation for closed-loop policy learning, and efficient inference. Taken together, this survey aims to provide a concise technical foundation for integrating predictive world modeling with action generation, toward more reliable embodied robot intelligence. Project page: https://rcl-robotics.github.io/Awesome-World-Action-Models.

Comments19 pages

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑