arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2511.01224cs.RO

视觉-语言-动作模型的具身迁移学习

Embodiment Transfer Learning for Vision-Language-Action Models

发表机构上海大学
查看机构详情
  • Shanghai University(上海大学)

机构由 AI 辅助整理,请以论文原文为准。

Chengmeng Li, Yaxin Peng

首次发表 更新
浏览论文内容

中文总结 AI 辅助

提出ET-VLA框架,通过合成持续预训练(SCP)和具身思维图技术,高效将预训练视觉-语言-动作模型迁移至多机器人协作,在真实任务上显著超越OpenVLA。

中文摘要 AI 辅助

视觉-语言-动作(VLA)模型显著推进了机器人学习,使其能够在大规模跨具身数据上训练并针对特定机器人进行微调。然而,当前最先进的自回归VLA在多机器人协作方面仍存在困难。我们引入了具身迁移学习(ET-VLA),这是一种用于将预训练VLA高效且有效迁移至多机器人的新框架。ET-VLA的核心是合成持续预训练(SCP),它利用合成生成的数据为新具身预热模型,从而无需真实的人类演示并降低了数据收集成本。SCP使模型能够学习正确的动作和精确的动作token数量。在SCP之后,模型会在目标具身数据上进行微调。为进一步提升模型在多具身上的性能,我们提出了具身思维图(Embodied Graph-of-Thought)技术,这是一种将每个子任务构建为节点的新方法,使得VLA模型能够在任务执行期间区分每种具身的功能和角色。我们的工作考虑了双臂机器人这一多机器人的简单版本来验证我们的方法。我们在覆盖三种不同双臂具身的模拟基准和真实机器人上验证了该方法的有效性。特别是,我们提出的ET-VLA在六项真实世界任务上的表现比OpenVLA高出53.2%以上。我们将开源所有代码,以支持社区推进用于机器人学习的VLA模型。

英文摘要

Vision-language-action (VLA) models have significantly advanced robotic learning, enabling training on large-scale, cross-embodiment data and fine-tuning for specific robots. However, state-of-the-art autoregressive VLAs struggle with multi-robot collaboration. We introduce embodiment transfer learning, denoted as ET-VLA, a novel framework for efficient and effective transfer of pre-trained VLAs to multi-robot. ET-VLA's core is Synthetic Continued Pretraining (SCP), which uses synthetically generated data to warm up the model for the new embodiment, bypassing the need for real human demonstrations and reducing data collection costs. SCP enables the model to learn correct actions and precise action token numbers. Following SCP, the model is fine-tuned on target embodiment data. To further enhance the model performance on multi-embodiment, we present the Embodied Graph-of-Thought technique, a novel approach that formulates each sub-task as a node, that allows the VLA model to distinguish the functionalities and roles of each embodiment during task execution. Our work considers bimanual robots, a simple version of multi-robot to verify our approaches. We validate the effectiveness of our method on both simulation benchmarks and real robots covering three different bimanual embodiments. In particular, our proposed ET-VLA \space can outperform OpenVLA on six real-world tasks over 53.2%. We will open-source all codes to support the community in advancing VLA models for robot learning.

↑