发表机构
Inria; PSL Research University(法国国家信息与自动化研究所; 巴黎文理研究大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出SA-WAM,将预训练视频模型用于联合动作、RGB与深度预测,在RoboCasa等基准及UR5机械臂真实评估中取得最优结果,为WAM改进提供了见解。
AI 中文摘要
世界动作模型(World Action Models,WAMs)利用大规模预训练视频扩散模型的能力,联合预测未来观测与动作,继承了互联网级视频蕴含的丰富视觉与物理先验,已成为机器人策略学习的有前景范式。但现有主流模型仅基于RGB观测,未利用3D信息。为填补该缺口,我们提出空间感知世界动作模型(Spatially Aware World Action Model,SA-WAM),将预训练视频模型重新用于联合动作、RGB与深度预测,在单个扩散主干中实现3D感知的世界建模与动作预测。我们采用非线性编码,将无界深度信号映射至冻结VAE分词器所需的有界输入域,从而可复用该分词器而无需3D特定微调,在不损失预训练先验的前提下融入几何信息。SA-WAM在RoboCasa与LIBERO-Plus基准上达到了当前最优结果,同时改进了未来状态预测;此外,在使用UR5机械臂的真实世界评估中,SA-WAM优于强基线方法,在随机环境中取得显著提升。我们分析了世界模型预测质量与回滚成功的相关性,为WAM的性能及改进方向提供了见解。
英文摘要
World Action Models (WAMs) leverage the capabilities of large-scale pretrained video diffusion models to jointly predict future observations and actions, inheriting rich visual and physical priors from internet-scale video. This has made them a promising paradigm for robot policy learning, yet the prevailing models operate exclusively on RGB observations and do not leverage 3D information. To bridge this gap, we introduce a Spatially Aware World Action Model (SA-WAM), which repurposes a pretrained video model for joint action, RGB, and depth prediction, enabling 3D-aware world modeling and action prediction within a single diffusion backbone. We use a nonlinear encoding that maps the unbounded depth signal into the bounded input domain expected by the frozen VAE tokenizer. This allows us to reuse the tokenizer without 3D-specific fine-tuning, incorporating geometric information without sacrificing the pretrained priors. SA-WAM achieves state-of-the-art results on the RoboCasa and LIBERO-Plus benchmarks, while simultaneously improving future-state predictions. Furthermore, SA-WAM outperforms strong baselines in real-world evaluation using a UR5 robotic arm, with strong gains in randomized environments. We analyze the correlation between world model prediction quality and rollout success, providing insights into WAM performance and avenues for its improvement.