arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2502.05855cs.ROcs.CV

DexVLA:带有插即用扩散专家的视觉-语言模型用于通用机器人控制

DexVLA: Vision-Language Model with Plug-In Diffusion Expert for General Robot Control

Junjie Wen, Yichen Zhu, Jinming Li, Zhibin Tang, Chaomin Shen, Feifei Feng

更新

AI总结:

本文提出DexVLA框架,通过引入十亿参数规模的插即用扩散动作专家和本体课程学习策略,解决VLA模型动作空间表示瓶颈,实现跨本体的高效训练与泛化,在单臂、双臂及灵巧手等复杂长时程任务上优于现有模型。

AI中文摘要:

使机器人能够在不同环境中执行多样化任务是机器人学习中的一个核心挑战。虽然视觉-语言-动作(VLA)模型在可泛化的机器人技能方面展现了潜力,但要充分发挥其潜力,需要解决动作表示和高效训练方面的局限性。当前的VLA模型通常侧重于扩展视觉-语言模型(VLM)组件,而动作空间表示仍然是一个关键瓶颈。本文介绍了DexVLA,这是一个新颖的框架,旨在增强VLA在跨多种机器人本体执行复杂、长时程任务时的效率和泛化能力。DexVLA采用了一个新颖的基于扩散的动作专家,参数规模达到十亿,专为跨本体学习设计。一种新颖的本体课程学习策略促进了高效训练:(1)在与VLA分离的跨本体数据上预训练扩散专家,(2)将VLA模型与特定本体对齐,(3)进行后训练以快速适应新任务。我们在多种本体上进行了全面实验,包括单臂、双臂和灵巧手,证明DexVLA无需特定任务适应即可适应挑战性任务,能够在有限数据下学习新本体的灵巧技能,并且仅通过直接语言提示就能完成复杂的、长时程任务,例如叠衣服。在所有设置中,我们的方法与Octo、OpenVLA和Diffusion Policy等最先进模型相比,表现出更优越的性能。

英文摘要:

Enabling robots to perform diverse tasks across varied environments is a central challenge in robot learning. While vision-language-action (VLA) models have shown promise for generalizable robot skills, realizing their full potential requires addressing limitations in action representation and efficient training. Current VLA models often focus on scaling the vision-language model (VLM) component, while the action space representation remains a critical bottleneck. This paper introduces DexVLA, a novel framework designed to enhance the efficiency and generalization capabilities of VLAs for complex, long-horizon tasks across diverse robot embodiments. DexVLA features a novel diffusion-based action expert, scaled to one billion parameters, designed for cross-embodiment learning. A novel embodiment curriculum learning strategy facilitates efficient training: (1) pre-training the diffusion expert that is separable from the VLA on cross-embodiment data, (2) aligning the VLA model to specific embodiments, and (3) post-training for rapid adaptation to new tasks. We conduct comprehensive experiments across multiple embodiments, including single-arm, bimanual, and dexterous hand, demonstrating DexVLA's adaptability to challenging tasks without task-specific adaptation, its ability to learn dexterous skills on novel embodiments with limited data, and its capacity to complete complex, long-horizon tasks using only direct language prompting, such as laundry folding. In all settings, our method demonstrates superior performance compared to state-of-the-art models like Octo, OpenVLA, and Diffusion Policy.

补充信息

↑