通过基于流(Flow)的视频预测迈向可泛化的双臂基础策略
Towards a Generalizable Bimanual Foundation Policy via Flow-based Video Prediction
- Institute of Artificial intelligence (TeleAI), China Telecom(人工智能研究院(TeleAI),中国电信)
- Northwestern Polytechnical University(西北工业大学)
- Hong Kong University of Science and Technology(香港科技大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对双臂操作数据稀缺与泛化难的问题,提出基于光流的两阶段视频预测范式微调文本到视频模型,结合轻量级扩散策略生成动作,有效降低数据需求并提升双臂操作泛化能力。
AI中文摘要:
由于动作空间庞大且需要协调手臂运动,学习可泛化的双臂操作策略对具身智能体而言极具挑战。现有方法依赖 Vision-Language-Action (VLA) 模型来获取双臂策略。然而,从单臂数据集或预训练 VLA 模型迁移知识往往无法有效泛化,主要原因是双臂数据稀缺以及单臂与双臂操作存在根本差异。本文提出一种新型双臂基础策略,通过微调领先的文本到视频模型来预测机器人轨迹,并训练轻量级扩散策略以生成动作。鉴于文本到视频模型缺乏具身知识,我们引入两阶段范式,微调源自预训练文本到视频模型的独立的 text-to-flow 和 flow-to-video 模型。具体而言,光流作为中间变量,提供了图像间细微运动的简洁表示。text-to-flow 模型预测光流以具体化语言指令的意图,flow-to-video 模型利用该光流进行细粒度视频预测。该方法缓解了单阶段文本到视频预测中语言的歧义,并通过避免直接使用低级动作显著降低了对机器人数据的需求。实验中,我们为真实双臂机器人收集了高质量操作数据,仿真和真实世界实验结果证明了该方法的有效性。
英文摘要:
Learning a generalizable bimanual manipulation policy is extremely challenging for embodied agents due to the large action space and the need for coordinated arm movements. Existing approaches rely on Vision-Language-Action (VLA) models to acquire bimanual policies. However, transferring knowledge from single-arm datasets or pre-trained VLA models often fails to generalize effectively, primarily due to the scarcity of bimanual data and the fundamental differences between single-arm and bimanual manipulation. In this paper, we propose a novel bimanual foundation policy by fine-tuning the leading text-to-video models to predict robot trajectories and training a lightweight diffusion policy for action generation. Given the lack of embodied knowledge in text-to-video models, we introduce a two-stage paradigm that fine-tunes independent text-to-flow and flow-to-video models derived from a pre-trained text-to-video model. Specifically, optical flow serves as an intermediate variable, providing a concise representation of subtle movements between images. The text-to-flow model predicts optical flow to concretize the intent of language instructions, and the flow-to-video model leverages this flow for fine-grained video prediction. Our method mitigates the ambiguity of language in single-stage text-to-video prediction and significantly reduces the robot-data requirement by avoiding direct use of low-level actions. In experiments, we collect high-quality manipulation data for real dual-arm robot, and the results of simulation and real-world experiments demonstrate the effectiveness of our method.