arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2509.00576cs.ROcs.CV

Galaxea开放世界数据集与G0双系统VLA模型

Galaxea Open-World Dataset and G0 Dual-System VLA Model

  • Galaxea Team(Galaxea 团队)

机构由 AI 辅助整理,请以论文原文为准。

Tao Jiang, Tianyuan Yuan, Yicheng Liu, Chenhao Lu, Jianning Cui, Xiao Liu, Shuiqi Cheng, Jiyang Gao, Huazhe Xu, Hang Zhao

更新

AI总结:

本研究提出真实环境采集的Galaxea开放世界机器人数据集,并构建结合VLM规划与VLA执行的G0双系统框架,经三阶段训练后在多项操作基准中表现优异。

AI中文摘要:

我们提出Galaxea Open-World Dataset(Galaxea开放世界数据集),这是在真实人类生活与工作环境中记录的大规模、多样化机器人行为集合。所有演示数据均采用统一的机器人实体采集,并搭配精确的子任务级语言标注,以同时支持训练与评估。基于该数据集,我们推出G0——一种双系统框架,将用于多模态规划的Vision-Language Model(VLM,视觉语言模型)与用于细粒度执行的Vision-Language-Action(VLA,视觉语言动作)模型相结合。G0采用三阶段课程训练:跨实体预训练、单实体预训练以及任务特定后训练。一项涵盖桌面操作、少样本学习和长时序移动操作的综合基准测试证明了我们方法的有效性。我们特别发现,单实体预训练阶段与Galaxea Open-World Dataset共同对实现优异性能起到了关键作用。

英文摘要:

We present Galaxea Open-World Dataset, a large-scale, diverse collection of robot behaviors recorded in authentic human living and working environments. All demonstrations are gathered using a consistent robotic embodiment, paired with precise subtask-level language annotations to facilitate both training and evaluation. Building on this dataset, we introduce G0, a dual-system framework that couples a Vision-Language Model (VLM) for multimodal planning with a Vision-Language-Action (VLA) model for fine-grained execution. G0 is trained using a three-stage curriculum: cross-embodiment pre-training, single-embodiment pre-training, and task-specific post-training. A comprehensive benchmark spanning tabletop manipulation, few-shot learning, and long-horizon mobile manipulation, demonstrates the effectiveness of our approach. In particular, we find that the single-embodiment pre-training stage, together with the Galaxea Open-World Dataset, plays a critical role in achieving strong performance.

补充信息

↑