arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2406.16862cs.ROcs.CV

Dreamitate:通过视频生成进行真实世界视觉运动策略学习

Dreamitate: Real-World Visuomotor Policy Learning via Video Generation

Junbang Liang, Ruoshi Liu, Ege Ozguroglu, Sruthi Sudhakar, Achal Dave, Pavel Tokmakov, Shuran Song, Carl Vondrick

首次发表 更新
浏览论文内容

中文总结 AI 辅助

针对机器人操作策略泛化难题,提出Dreamitate框架,通过在人类演示上微调预训练视频扩散模型,以生成的任务执行视频直接控制机器人,在四项任务上实现了比现有行为克隆方法更优的泛化性能。

中文摘要 AI 辅助

操作任务中的一项核心挑战是学习能够稳健泛化到多样视觉环境的策略。学习鲁棒策略的一种可行机制是利用在大规模互联网视频数据集上预训练的视频生成模型。本文提出一种视觉运动策略学习框架Dreamitate,针对给定任务的人类演示数据对视频扩散模型进行微调。测试阶段,以新场景的图像为条件生成任务执行示例,直接使用该合成执行过程控制机器人。我们的核心洞见是,借助通用工具可轻松弥合人手与机器人机械臂之间的具身差异。我们在四项复杂度递增的任务上评估了所提方法,结果表明,利用互联网规模的生成模型,所学策略的泛化能力显著优于现有行为克隆方法。

英文摘要

A key challenge in manipulation is learning a policy that can robustly generalize to diverse visual environments. A promising mechanism for learning robust policies is to leverage video generative models, which are pretrained on large-scale datasets of internet videos. In this paper, we propose a visuomotor policy learning framework that fine-tunes a video diffusion model on human demonstrations of a given task. At test time, we generate an example of an execution of the task conditioned on images of a novel scene, and use this synthesized execution directly to control the robot. Our key insight is that using common tools allows us to effortlessly bridge the embodiment gap between the human hand and the robot manipulator. We evaluate our approach on four tasks of increasing complexity and demonstrate that harnessing internet-scale generative models allows the learned policy to achieve a significantly higher degree of generalization than existing behavior cloning approaches.

发表机构

  • Columbia University(哥伦比亚大学)
  • Toyota Research Institute(丰田研究所)
  • Stanford University(斯坦福大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑