发表机构
Columbia University; Toyota Research Institute(哥伦比亚大学; 丰田研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
SUAVE提出单一词汇表的掩码扩散Transformer,将视频和动作统一为离散token,实现世界模型、策略和视频-动作模型三合一,并在仿真和真实机器人上验证了性能与鲁棒性。
AI 中文摘要
视觉-语言-动作模型(VLAs)继承了预训练视觉-语言骨干网络的强大语义基础,但通常针对预测动作而非未来观测进行优化。它们能够观察和行动,但在行动之前不会想象未来。基于视频扩散骨干网络构建的世界动作模型(WAMs)能够想象,但将语言视为连续潜在空间上的冻结条件。统一模型将这些模态整合到一个架构中,但它们要么以自回归方式逐token解码,要么通过辅助动作头保持视频连续性。在这项工作中,我们提出了SUAVE,一种单一词汇表的统一动作-视频模型,其中掩码扩散Transformer在语言条件下生成视频和动作,所有三种模态在共享序列中表示为离散token。在推理时选择要掩码的token,可将同一网络转变为世界模型、机器人策略或视频-动作模型。对于无动作的协同训练,未标记视频的动作位置填充掩码token并从损失中排除。仿真和真实世界实验展示了两个发现。首先,单个SUAVE模型预测长时程视频并充当策略,在静态和动态操作任务上与专用世界模型和专门动作策略具有竞争力。在真实机器人上,我们的模型在RTX 5090 GPU上以1,030毫秒生成子目标图像和覆盖一秒运动范围的动作块,以每秒2.5个动作维持闭环控制。其次,在机器人视频上的预训练和在人类视频上的协同训练显著提高了策略性能和面对分布偏移的零样本鲁棒性。这些结果共同表明,掩码扩散是统一视频-动作建模的一种实用且多功能的基石。
英文摘要
Vision-language-action models (VLAs) inherit strong semantic grounding from pretrained vision-language backbones but are typically optimized for predicting actions rather than future observations. They can see and act, but they do not imagine the future before acting. World action models (WAMs) built on video diffusion backbones can imagine but treat language as frozen conditioning on a continuous latent space. Unified models bring these modalities into one architecture, but they either decode autoregressively, one token at a time, or keep video continuous with an auxiliary action head. In this work, we present SUAVE, a Single vocabulary Unified Action-Video modEl in which a masked diffusion transformer generates video and actions conditioned on language, with all three modalities represented as discrete tokens in a shared sequence. Choosing which tokens to mask at inference turns the same network into a world model, a robot policy, or a video-action model. For action-free co-training, the action positions of unlabeled video are filled with mask tokens and excluded from the loss. Simulation and real-world experiments demonstrate two findings. First, a single SUAVE model predicts long-horizon video and acts as a policy, competitive with dedicated world models and specialized action policies on static and dynamic manipulation tasks. On a real robot, our model generates subgoal images and an action chunk spanning one second of motion in 1,030 ms on an RTX 5090 GPU, sustaining closed-loop control at 2.5 actions per second. Second, pretraining on robot video and co-training on human video substantially improves policy performance and zero-shot robustness to distribution shift. Together, these results show that masked diffusion is a practical and versatile foundation for unified video-action modeling.
CommentsPreprint version