发表机构
The Hong Kong University of Science and Technology (Guangzhou); Beihang University; Huawei Foundation Model Department(香港科技大学(广州); 北京航空航天大学; 华为基础模型部门)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
Robust-WAM是一种通用后训练方法,在保留VAE生成路径的同时添加语义前瞻对齐,可提升多个WAM基线在分布外条件下的动作预测成功率且不损失分布内性能。
AI 中文摘要
主流世界-动作模型(World-Action Models, WAMs)将预训练视频生成模型(Video Generation Models, VGMs)适配于机器人控制,迁移其学习到的动态先验以进行动作预测。这些VGMs通常在变分自编码器(Variational Autoencoder, VAE)的隐空间中训练,然而该隐空间以像素重建为优化目标,侧重精细外观细节,导致动作预测在视觉偏移下较为脆弱。近期研究在语义隐空间中构建WAMs,这类模型对外观偏移更具鲁棒性,但无法利用仅存在于VAE空间的大规模VGM预训练。为克服该困境,我们提出Robust-WAM,一种基于视频生成的WAMs的通用后训练方法,其保留基于VAE的生成路径,并在动作流中添加轻量型语义前瞻对齐目标。这既保留了大规模VGM预训练,又将动作锚定在外观不变的动态中,该动态在光照偏移及其他视觉分布外条件下仍保持可靠。具体而言,我们采用可学习查询令牌,通过将其输出隐藏状态与未来真实帧的语义前瞻对齐,将未来场景语义引入动作流;为建立每个查询与其描述的未来步之间的时间对应关系,我们为其赋予匹配动作令牌的位置编码。在分布外泛化仿真基准及真实机器人设置上的实验表明,Robust-WAM在不损失分布内性能的前提下,持续提升了多个WAM基线的成功率。
英文摘要
Mainstream World-Action Models (WAMs) adapt pretrained video generation models (VGMs) for robot control, transferring their learned dynamics prior for action prediction. These VGMs are typically trained in a variational autoencoder (VAE) latent space. However, the VAE latent space is optimized for pixel reconstruction, which rewards fine appearance detail and leaves the action prediction fragile under visual shifts. Recent works build WAMs in semantic latent space, which are more robust to appearance shifts. However, these models cannot leverage the large-scale VGM pretraining that exists only in VAE space. To overcome this dilemma, we propose Robust-WAM, a general post-training method for video-generation-based WAMs that preserves the VAE-based generative path and adds a lightweight semantic foresight alignment objective on the action stream. This retains the large-scale VGM pretraining while grounding actions in appearance-invariant dynamics that stay reliable under illumination shifts and other visual out-of-distribution conditions. Specifically, we employ learnable query tokens to bring future-scene semantics into the action stream by aligning their output hidden states with the semantic foresight of future ground-truth frames. To establish the temporal correspondence between each query and the future step it describes, we give it the positional encoding of the matching action tokens. Experiments on out-of-distribution generalization simulation benchmarks and a real-robot setup show that our Robust-WAM consistently improves the success rates of multiple WAM baselines without sacrificing in-distribution performance.