arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2602.22663cs.RO

重新审视视觉-语言-动作模型的实用性:一个全面的基准和一个改进的基线

Rethinking the Practicality of Vision-language-action Model: A Comprehensive Benchmark and An Improved Baseline

Wenxuan Song, Jiayi Chen, Xiaoquan Sun, Huashuo Lei, Yikai Qin, Wei Zhao, Pengxiang Ding, Han Zhao, Tongxin Wang, Pengxu Hou, Zhide Zhong, Haodong Yan, Donglin … 展开作者

Wenxuan Song, Jiayi Chen, Xiaoquan Sun, Huashuo Lei, Yikai Qin, Wei Zhao, Pengxiang Ding, Han Zhao, Tongxin Wang, Pengxu Hou, Zhide Zhong, Haodong Yan, Donglin Wang, Jun Ma, Haoang Li

首次发表 更新
浏览论文内容

中文总结 AI 辅助

本文提出CEbench基准和LLaVA-VLA模型,通过轻量级设计和两阶段训练提升视觉-语言-动作模型的实用性,实现在消费级GPU上的实际应用。

中文摘要 AI 辅助

视觉-语言-动作(VLA)模型已经作为一种通用的机器人代理出现。然而,现有的VLA受限于过大的参数规模、昂贵的预训练要求以及对多样化实体应用的限制。为了提高VLA的实用性,我们提出了一个全面的基准和一个改进的基线。首先,我们提出了CEbench,一个新的基准,涵盖了模拟和现实世界中的多样化实体,并考虑了领域随机化。我们收集了14400个模拟轨迹和1600个现实世界专家整理的轨迹,以支持在CEbench上进行训练。其次,利用CEbench作为我们的测试平台,我们研究了VLA实用性三个关键方面,并提供了几个关键发现。受这些发现的启发,我们引入了LLaVA-VLA,一种轻量但强大的VLA,专为在消费级GPU上进行实际部署设计。在结构上,它集成了紧凑的VLM主干网络,多视角感知,本体编码和动作分块。为了消除对昂贵预训练的依赖,LLaVA-VLA采用了两阶段训练范式,包括预训练后的训练和微调。此外,LLaVA-VLA扩展了动作空间,以统一导航和操作。在不同实体上的实验展示了LLaVA-VLA的泛化能力和多功能性,而现实世界的移动操作实验将其确立为首个端到端的VLA模型用于移动操作。我们将在接受后开源所有数据集、代码和检查点,以促进可重复性和未来研究。

英文摘要

Vision-Language-Action (VLA) models have emerged as a generalist robotic agent. However, existing VLAs are hindered by excessive parameter scales, prohibitive pre-training requirements, and limited applicability to diverse embodiments. To improve the practicality of VLAs, we propose a comprehensive benchmark and an improved baseline. First, we propose CEBench, a new benchmark spanning diverse embodiments in both simulation and the real world with consideration of domain randomization. We collect 14.4k simulated trajectories and 1.6k real-world expert-curated trajectories to support training on CEBench. Second, using CEBench as our testbed, we study three critical aspects of VLAs' practicality and offer several key findings. Informed by these findings, we introduce LLaVA-VLA, a lightweight yet powerful VLA designed for practical deployment on consumer-grade GPUs. Architecturally, it integrates a compact VLM backbone with multi-view perception, proprioceptive tokenization, and action chunking. To eliminate reliance on costly pre-training, LLaVA-VLA adopts a two-stage training paradigm including post-training and fine-tuning. Furthermore, LLaVA-VLA extends the action space to unify navigation and manipulation. Experiments across embodiments demonstrate the capabilities of generalization and versatility of LLaVA-VLA , while real-world mobile manipulation experiments establish it as the first end-to-end VLA model for mobile manipulation. We will open-source all datasets, codes, and checkpoints upon acceptance to foster reproducibility and future research.

发表机构

  • OpenHelix-Team(OpenHelix团队)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑