通过大规模无动作视频预训练学习可执行的离散扩散策略
Learning an Actionable Discrete Diffusion Policy via Large-Scale Actionless Video Pre-Training
浏览论文内容
中文总结 AI 辅助
针对带动作标签的机器人数据集稀缺问题,提出结合人类视频生成式预训练与少量机器人数据策略微调的统一离散扩散框架,可生成高保真未来视频并提升机器人策略性能。
中文摘要 AI 辅助
学习一个能够完成多项任务的通用具身智能体存在诸多挑战,这些挑战主要源于带动作标签的机器人数据集稀缺。相比之下,现有大量人类视频记录了复杂任务以及人类与物理世界的交互过程。利用无动作的人类视频进行预训练,并迁移相关知识以通过少量机器人演示辅助机器人策略学习,具备良好的应用前景。但由于人类与机器人之间存在领域差异,这一目标仍面临挑战。此外,人类视频的数据结构存在噪声且具有多模态特性,很难从中提取出能够表征动态世界的有效信息。本文提出了一个全新框架以应对上述挑战,该框架借助统一离散扩散模型,将人类视频上的生成式预训练与少量带动作标签的机器人视频上的策略微调相结合。我们首先将人类视频和机器人视频都压缩为统一的视频令牌(token)。在预训练阶段,我们采用带有掩码-替换扩散策略的离散扩散模型,在隐空间中预测未来的视频令牌。在微调阶段,我们利用生成的想象未来视频,在少量机器人数据的辅助下指导底层动作学习。实验表明,与现有最优方法相比,本文方法能够生成用于规划的高保真未来视频,且提升了微调后策略的性能,表现更优。我们的项目网站为:https://video-diff.github.io/。
英文摘要
Learning a generalist embodied agent capable of completing multiple tasks poses challenges, primarily stemming from the scarcity of action-labeled robotic datasets. In contrast, a vast amount of human videos exist, capturing intricate tasks and interactions with the physical world. Promising prospects arise for utilizing actionless human videos for pre-training and transferring the knowledge to facilitate robot policy learning through limited robot demonstrations. However, it remains a challenge due to the domain gap between humans and robots. Moreover, it is difficult to extract useful information representing the dynamic world from human videos, because of its noisy and multimodal data structure. In this paper, we introduce a novel framework to tackle these challenges, which leverages a unified discrete diffusion to combine generative pre-training on human videos and policy fine-tuning on a small number of action-labeled robot videos. We start by compressing both human and robot videos into unified video tokens. In the pre-training stage, we employ a discrete diffusion model with a mask-and-replace diffusion strategy to predict future video tokens in the latent space. In the fine-tuning stage, we harness the imagined future videos to guide low-level action learning with a limited set of robot data. Experiments demonstrate that our method generates high-fidelity future videos for planning and enhances the fine-tuned policies compared to previous state-of-the-art approaches with superior performance. Our project website is available at https://video-diff.github.io/.
发表机构
- Hong Kong University of Science and Technology(香港科技大学)
- Institute of Artificial Intelligence (TeleAI), China Telecom(中国电信人工智能研究院(TeleAI))
- Shanghai Jiao Tong University(上海交通大学)
- Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)
机构由 AI 辅助整理,请以论文原文为准。