arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

OpenWAM:面向系统性世界-动作模型预训练的开源模块化探索

OpenWAM: An Open, Modular Exploration Towards Systematic World-Action Model Pretraining

Yuran Wang, Siqiao Huang, Mingleyang Li, Chenhao Zhang, Jiaqi Liang, Weiyang Jin, Yue Chen, Xuemin Chi, Donghao Zhou, Qize Yu, Yu-Kai Wang, Yuhan Rui, Shenzhe Yao, Zhen Yuan, Zhenhao Shen, Kefei Zhu, Zijie Zhu, Ning Gao, Xiaowei Chi, Guanqi He, Shanghang Zhang, Hao Dong, Lin Shao, Hang Zhao

arXiv 2609.07398首次发表:更新:

发表机构

National University of Singapore; Tsinghua University; Peking University; The University of Hong Kong; Zhejiang University; The Chinese University of Hong Kong; Shanghai Jiao Tong University(新加坡国立大学; 清华大学; 北京大学; 香港大学; 浙江大学; 香港中文大学; 上海交通大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

OpenWAM提出开源模块化框架,将世界-动作预训练分解为可控组件,通过受控实验提炼设计原则,并基于约6400小时数据预训练出在仿真和真实世界基准上表现优异的模型。

AI 中文摘要

世界-动作模型从视频生成先验中继承世界知识,并通过具身经验将其转化为可执行的控制信号。然而,现有系统是整体式的:生成骨干网络、视觉表示、架构、信息流、推理过程和训练数据紧密耦合,模糊了哪些设计选择至关重要及其原因。我们提出OpenWAM,一个开源研究栈,将世界-动作预训练转变为受控的实验项目。OpenWAM-Infra将WAM设计空间分解为可组合的模块,并统一训练、推理、部署和评估。在此基础之上,OpenWAM-Study通过受控实验探讨三个问题:继承什么、世界学习与动作学习如何交互、以及它们的协同如何扩展;并提炼出三项原则:上游知识通过足够强大的生成骨干网络和紧凑、信息丰富的潜在空间进行迁移;世界-动作协同需要专用的动作容量、显式的世界到动作信息流以及同步的联合去噪;具身预训练主要改善域外泛化,其中基于自我中心数据和机器人数据的一阶段协同训练整合了世界覆盖和动作基础。综合这些原则,我们构建了OpenWAM-α,一个在约6,400小时的自我中心人类和机器人数据上预训练的开源WAM,并在仿真和真实世界基准上进行了评估。在跨越单臂、双臂操作到灵巧手等具身形态的八个仿真基准和真实机器人实验中,OpenWAM-α始终表现出色,从仿真到物理世界保持了顶级水平。我们发布了完整的技术栈,包括基础设施、评估协议、预训练模型和数据配方,以促进未来研究。

英文摘要

World-Action Models inherit world knowledge from video-generative priors, and channel it into executable control signals through embodied experience. Existing systems, however, are monolithic: the generative backbone, visual representation, architecture, information flow, inference procedure, and training data are tightly coupled, obscuring which design choices matter and why. We introduce OpenWAM, an open research stack that turns world-action pretraining into a controlled experimental program. OpenWAM-Infra factorizes the WAM design space into composable modules with unified training, inference, deployment, and evaluation. On this substrate, OpenWAM-Study examines three questions through controlled experiments: what to inherit, how world and action learning interact, and how their synergy scales; and distills three principles: upstream knowledge transfers through a sufficiently capable generative backbone and a compact, information-rich latent space; world-action synergy requires dedicated action capacity, explicit world-to-action information flow, and synchronized joint denoising; and embodied pretraining principally improves out-of-domain generalization, with one-stage co-training over egocentric and robot data integrating world coverage and action grounding. Composing these principles, we build OpenWAM-α, an open WAM pretrained on roughly 6,400 hours of egocentric human and robot data and evaluated across simulation and real-world benchmarks. Across the eight simulation benchmarks and the real-robot experiments, which together span embodiments from single-arm and bimanual manipulation to dexterous hands, OpenWAM-α delivers consistently excellent performance, sustaining its top-tier standing from simulation to the physical world. We release the full stack, including infrastructure, evaluation protocols, pretrained models, and data recipes, to facilitate future research.

CommentsProject Page: https://openwam-official.github.io/; Code: https://github.com/OpenWAM-Official/OpenWAM; Model & Data: https://huggingface.co/OpenWAM

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑