arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

设计世界动作模型的关键因素:一项实证研究

What Matters in Designing World Action Models: An Empirical Study

Chao Tang, Haoqing Wang, Zilang Cen, Weishi Mi, Wei Xia, Fangcheng Liu, Anda Cheng, Yeqing Shen, Xiaohui Cui, Xiaoyuan Zhang, Yehui Tang, Tingguang Li

arXiv 2609.24048首次发表:更新:

发表机构

Samsung Robotics eXperience; Samsung R&D Institute China–Beijing; Zhongguancun Academy; Wuhan University; Peking University(三星机器人体验中心; 三星电子中国研究院北京分院; 中关村学院; 武汉大学; 北京大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文通过受控实验系统比较世界动作模型中的因果结构、潜在空间和训练目标,揭示核心设计选择的影响,为通用机器人控制提供指导原则。

AI 中文摘要

世界动作模型(WAMs)已成为通用机器人控制的一种有前景的范式。尽管WAM系统的数量不断增长,现有工作往往引入将多种设计选择(如架构和训练策略)捆绑在一起的统一系统,这使得难以隔离各个贡献并系统地比较替代设计。在这项工作中,我们提出了一项受控研究,以解耦这些设计选择,并不仅分析其经验效果,还分析它们如何以及为何塑造WAMs。更具体地,我们聚焦于构建WAMs中的三个基本问题:(1)世界建模与动作生成之间的交互应由何种因果结构支配?(2)世界建模应在何种潜在空间中进行?(3)不同的世界-动作建模目标如何影响模型行为和性能?通过在三个代表性基准(RoboCasa-GR1、LIBERO和LIBERO-Plus)上进行结构受控实验,我们系统地比较了六种因果结构、八种潜在表示和四种训练目标,涵盖了现有WAMs中的流行设计选择。我们进一步在来自DROID数据集的真实机器人数据上验证了我们的关键发现。我们希望提供对核心设计选择如何影响世界-动作建模以及哪些原则可以指导未来WAM系统开发的系统性理解。

英文摘要

World Action Models (WAMs) have emerged as a promising paradigm for generalizable robot control. Despite the growing number of WAM systems, existing works often introduce unified systems that bundle together multiple design choices, such as architecture and training strategy, making it difficult to isolate individual contributions and systematically compare alternative designs. In this work, we present a controlled study that disentangles these design choices and analyzes not only their empirical effects, but also how and why they shape WAMs. More specifically, we focus on three fundamental questions in building WAMs: (1) what causal structure should govern the interaction between world modeling and action generation? (2) in which latent space should world modeling be performed? and (3) how do different world-action modeling objectives affect model behavior and performance? Through structurally controlled experiments on three representative benchmarks, RoboCasa-GR1, LIBERO, and LIBERO-Plus, we systematically compare six causal structures, eight latent representations, and four training objectives, covering popular design choices in existing WAMs. We further validate our key findings on real-robot data from the DROID dataset. We hope to provide a systematic understanding of how core design choices affect world-action modeling and what principles can guide the development of future WAM systems.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑