arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

OpenWAM:一个用于可组合世界-动作模型的开放框架

OpenWAM: An Open Framework for Composable World-Action Models

Heng Yu, David D. Yuan, Juze Zhang, Changan Chen, Yao Feng, Michelle Baldonado, Steve Cousins, Li Fei-Fei, Jiajun Wu, Ehsan Adeli

arXiv 2610.07922首次发表:更新:

发表机构

Stanford University(斯坦福大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出开放框架OpenWAM,基于因果机器人-视频预训练和可配置交互架构,在LIBERO及真实任务上提升成功率,并利用反事实监督改善动力学模型学习。

AI 中文摘要

世界-动作模型(WAMs)将未来预测与机器人控制相结合,然而现有系统常常同时改变视频主干网络、交互结构、监督方式和推理过程,使得其设计选择难以比较。我们提出了OPENWAM,一个开放的世界-动作建模框架,它基于一个通用的因果机器人-视频基础模型和可配置的视频-动作交互。从Wan2.2-5B开始,我们在超过10,000小时的视频上进行因果机器人-视频预训练,然后通过一个共享的Mixture-of-Transformers架构集成一个动作专家,该架构支持联合生成、视频-然后-动作、动作-然后-视频以及解耦生成。OPENWAM在四个LIBERO套件和真实世界双臂任务上取得了高成功率;带有因果适配的机器人-视频训练将LIBERO-Long上的VTA成功率从68.4%提升至97.8%。相同的可配置架构自然扩展到逆动力学和正动力学,使我们能够研究反事实转换如何改善独立训练的动力学模型,而不仅仅依赖示范数据。当仅将视频预测器适配到新任务时,一个在反事实数据和示范上训练的冻结局部上下文逆动力学模型在四个保留的LIBERO-90任务上实现了84.0%的平均成功率,而全上下文逆模型为47.0%,仅使用示范训练的局部上下文模型为21.5%。对于正动力学,反事实监督将RGB预测误差降低了34.5%,并将16个相同状态结果中的结果识别率从21.1%提升至71.3%。OPENWAM为比较WAM交互设计以及研究从视频数据中学习动力学(超越成功示范)提供了一个通用测试平台。

英文摘要

World-action models (WAMs) couple future prediction with robot control, yet existing systems often vary the video backbone, interaction structure, supervision, and inference procedure simultaneously, making their design choices difficult to compare. We introduce OPENWAM, an open world-action modeling framework built around a common causal robot-video foundation and configurable video-action interaction. Starting from Wan2.2-5B, we perform causal robot-video pretraining on over 10,000 hours of video, then integrate an action expert through a shared Mixture-of-Transformers architecture that supports joint, video-then-action, action-then-video, and decoupled generation. OPENWAM achieves high success rates on four LIBERO suites and real-world bimanual tasks; robot-video training with causal adaptation improves VTA success on LIBERO-Long from 68.4% to 97.8%. The same configurable architecture naturally extends to inverse and forward dynamics, allowing us to study how counterfactual transitions improve independently trained dynamics models beyond demonstrations alone. When only the video predictor is adapted to a new task, a frozen local-context inverse dynamics model trained on counterfactual data and demonstrations achieves 84.0% mean success across four held-out LIBERO-90 tasks, compared with 47.0% for a full-context inverse model and 21.5% for a local-context model trained only on demonstrations. For forward dynamics, counterfactual supervision reduces RGB prediction error by 34.5% and raises outcome identification from 21.1% to 71.3% among 16 same-state outcomes. OPENWAM provides a common testbed for comparing WAM interaction designs and for studying dynamics learning from video data beyond successful demonstrations.

Comments18 pages, 5 figures, 14 tables. Project page: https://openwam.stanford.edu ; Code: https://github.com/OpenWAM/OpenWAM ; Code and project page released June 4, 2026. Equal contribution: Heng Yu, David D. Yuan, Juze Zhang

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

相关深度报道

↑