arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2412.14803cs.CVcs.RO

视频预测策略:一种具有预测性视觉表示的通用机器人策略

Video Prediction Policy: A Generalist Robot Policy with Predictive Visual Representations

Yucheng Hu, Yanjiang Guo, Pengchao Wang, Xiaoyu Chen, Yen-Jen Wang, Jianke Zhang, Koushil Sreenath, Chaochao Lu, Jianyu Chen

首次发表 更新
浏览论文内容

中文总结 AI 辅助

本文提出视频预测策略(VPP),利用视频扩散模型内部的预测性未来表示学习隐式逆动力学模型,在Calvin ABC-D基准上相对提升18.6%,并在真实世界灵巧操作中成功率提升31.6%。

中文摘要 AI 辅助

视觉表示在开发通用机器人策略中起着至关重要的作用。以往的视觉编码器通常通过单图像重建或双图像对比学习进行预训练,倾向于捕捉静态信息,往往忽视了对于具身任务至关重要的动态方面。近年来,视频扩散模型(VDMs)展示了预测未来帧的能力,并表现出对物理世界的强大理解。我们假设VDMs本质上产生的视觉表示同时包含当前静态信息和预测的未来动态,从而为机器人动作学习提供有价值的指导。基于这一假设,我们提出了视频预测策略(VPP),该策略在VDMs内部以预测的未来表示为条件学习隐式逆动力学模型。为了预测更精确的未来,我们在机器人数据集以及互联网人类操作数据上对预训练的视频基础模型进行微调。在实验中,与之前的最先进方法相比,VPP在Calvin ABC-D泛化基准上取得了18.6%的相对提升,并在复杂的真实世界灵巧操作任务中展示了31.6%的成功率提升。项目主页位于https://video-prediction-policy.github.io。

英文摘要

Visual representations play a crucial role in developing generalist robotic policies. Previous vision encoders, typically pre-trained with single-image reconstruction or two-image contrastive learning, tend to capture static information, often neglecting the dynamic aspects vital for embodied tasks. Recently, video diffusion models (VDMs) demonstrate the ability to predict future frames and showcase a strong understanding of physical world. We hypothesize that VDMs inherently produce visual representations that encompass both current static information and predicted future dynamics, thereby providing valuable guidance for robot action learning. Based on this hypothesis, we propose the Video Prediction Policy (VPP), which learns implicit inverse dynamics model conditioned on predicted future representations inside VDMs. To predict more precise future, we fine-tune pre-trained video foundation model on robot datasets along with internet human manipulation data. In experiments, VPP achieves a 18.6\% relative improvement on the Calvin ABC-D generalization benchmark compared to the previous state-of-the-art, and demonstrates a 31.6\% increase in success rates for complex real-world dexterous manipulation tasks. Project page at https://video-prediction-policy.github.io

发表机构

  • IIIS, Tsinghua University(清华大学信息科学技术学院)
  • Shanghai Qi Zhi Institute(上海期智研究院)
  • RobotEra

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑