EXIMO:基于视觉语言模型引导的视觉语言动作策略探索
EXIMO: VLM Guided Exploration of VLA Policies
查看机构详情
- Google DeepMind(谷歌DeepMind)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
本研究提出EXIMO算法,通过VLM引导的探索、模仿、优化三阶段微调VLA策略,解决了VLA策略微调样本效率低的问题,性能显著优于现有方法。
中文摘要 AI 辅助
如何高效微调机器人策略以实时学习新任务?当前最优的机器人操作策略基于在海量遥操作数据集上对拥有数十亿参数的大型视觉语言动作(VLA)模型进行行为克隆。尽管这种简单方法已推动机器人操作领域取得显著进展,但针对新任务微调VLA策略仍是一个未解决的问题。具体而言,收集遥操作数据集需要耗费数百小时的昂贵人力,而替代方案强化学习(RL)则以样本效率低著称,尤其对于长 horizon 任务。此外,由于VLA的模型规模和架构设计,使用VLA的RL还存在若干挑战。本研究提出EXIMO,一种用于微调VLA策略的高效算法,其分为三个阶段:探索、模仿与优化。在探索阶段,EXIMO为VLA配备作为规划器的视觉语言模型(VLM),该VLM将复杂的长 horizon 问题拆解为更短的子问题,VLM与VLA共同用于收集新任务的协调数据集;在模仿阶段,使用协调后的数据对VLA进行微调;最后在优化阶段,采用残差离策略RL进一步微调策略。实验中,我们对EXIMO的三个阶段进行了消融分析,结果显示其在样本效率和最终性能上显著优于现有方法。
英文摘要
How to efficiently finetune robot policies to learn new tasks on the fly? State of the art robotic manipulation policies are based on behaviour cloning of large vision-language-action (VLA) models with billions of parameters on huge teleoperation datasets. While this simple approach has enabled significant advances for robotic manipulation, finetuning of VLA policies for learning new tasks still remains an open problem. In particular, collecting teleoperation datasets requires hundreds of hours of expensive human labour and the alternative, reinforcement learning (RL), can be notoriously sample-inefficient especially for long-horizon tasks. In addition, RL with VLAs imposes several challenges due to the model's size and architectural design. In this work, we propose EXIMO, an efficient algorithm for finetuning of VLA policies. EXIMO operates in three stages: explore, imitate, and optimize. During the explore phase, EXIMO equips the VLA with a vision language model (VLM) that acts as a planner. The VLM thinks and breaks down challenging long-horizon problems into shorter ones for the VLA. The VLM, together with the VLA, is used to collect an orchestrated dataset on new tasks. During the imitate phase, the VLA is finetuned with the orchestrated data. Finally, during the optimize stage, we use residual off-policy RL to further finetune the policy. In our experiments, we ablate all three stages of EXIMO and show that it outperforms existing approaches significantly in terms of sample-efficiency and final performance.