视频预测策略2:预测更好,行动更好
Video Prediction Policy 2: Predict Better, Act Better
浏览论文内容
中文总结 AI 辅助
提出视频预测策略2(VPP2),通过大规模操作视频预训练、事件级预训练、蒸馏为单步规划器及混合变换器动作模块,在视频预测和零样本操作任务上显著超越现有方法。
中文摘要 AI 辅助
世界动作模型(WAMs)已成为通用机器人策略的重要类别,旨在将视频预测先验迁移到动作学习中。然而,我们发现现有的WAMs在开放环境中经常产生错误的运动预测,导致错误的动作。我们将此局限性归因于两个因素:(1)基础视频模型未针对操作进行优化,(2)将动作组件简单纳入视频模型会显著降低其泛化能力。我们引入了视频预测策略2(VPP2),一种在视频预测和动作生成中均能实现强大零样本泛化的WAM。首先,我们整理了一个大规模、多样化的操作视频数据集,以继续预训练基础视频基础模型。我们为视频片段标注了详细描述,并进行事件级视频预训练,以促进跨开放操作任务的泛化。其次,我们对视频模型进行后训练和蒸馏,将其转化为具有固定预测范围的单步视觉规划器。最后,我们通过混合变换器(MoT)架构引入动作模块,以学习隐式逆动力学模型。实验展示了三个关键结果:(1)VPP2-14B在开放任务上的视频预测指令遵循成功率上比Cosmos3-64B高出11.0个百分点;(2)VPP2在真实世界零样本ALOHA操作任务上的成功率比最强基线高出18.5个百分点;(3)在遵循基准特定后训练后,VPP2在具有挑战性的LIBERO-Pro、LIBERO-OOD和RoboDojo基准上取得了评估方法中最高的成功率。
英文摘要
World action models (WAMs) have emerged as an important class of generalist robot policies, aiming to transfer video prediction priors to action learning. However, we find that existing WAMs frequently produce incorrect motion predictions in open-ended environment, leading to erroneous actions. We attribute this limitation to two factors: (1) base video models are not optimized for manipulation, and (2) naively incorporating action components into video models can substantially degrade their generalization capabilities. We introduce Video Prediction Policy 2 (VPP2), a WAM that enables strong zero-shot generalization in both video prediction and action generation. First, we curate a large-scale, diverse dataset of manipulation videos to continue pretraining the base video foundation model. We annotate video clips with detailed captions and perform \textit{event-level} video pretraining to promote generalization across open-ended manipulation tasks. Second, we post-train and distill the video model into a single-step visual planner with fixed prediction horizon. Finally, we introduce action module via a mixture-of-transformers (MoT) architecture to learn implicit inverse dynamics model. Experiments demonstrate three key results: (1) VPP2-14B outperforms Cosmos3-64B by 11.0\% points in video prediction instruction-following success rate on open-ended tasks; (2) VPP2 surpasses the strongest baseline by 18.5\% points in success rate on real-world zero-shot ALOHA manipulation tasks; and (3) following benchmark-specific post-training, VPP2 achieves the highest success rates among evaluated methods on the challenging LIBERO-Pro, LIBERO-OOD, and RoboDojo benchmarks.
发表机构
- Robotera
- Tsinghua University(清华大学)
- HKUST (GZ)(香港科技大学(广州))
- University of California, Berkeley(加州大学伯克利分校)
- Shanghai Jiaotong University(上海交通大学)
机构由 AI 辅助整理,请以论文原文为准。