arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.11739cs.ROcs.AI

G0.5:用于机器人推理与动作的单自回归流

G0.5: One Autoregressive Stream for Robot Reasoning and Action

  • Galaxea(星系科技(Galaxea))

机构由 AI 辅助整理,请以论文原文为准。

Yicheng Liu, Zibin Dong, Baijun Ye, Tianyuan Yuan, Tao Jiang, Anqi Yang, Shicheng Cao, Haonan Liu, Yue Sun, Zihan Guo, Xiao Liu, Dong Ke, Changxun Pan, Chenru W… 展开作者

Yicheng Liu, Zibin Dong, Baijun Ye, Tianyuan Yuan, Tao Jiang, Anqi Yang, Shicheng Cao, Haonan Liu, Yue Sun, Zihan Guo, Xiao Liu, Dong Ke, Changxun Pan, Chenru Wu, Tailai Cheng, Xiaoshu Ren, Xinlei Zhang, Jianning Cui, Zijie Zhao, Haoyu Zhang, Kaiming Xu, Haodong Yang, Bowen Zhang, Jiahui Niu, Shaoting Zhu, Shiduo Zhang, Hang Zhao

AI总结:

本文提出G0.5,一种单自回归流的VLA模型,通过三个关键组件实现推理与动作统一,在7项基准中超越现有最优模型,展现出良好的泛化与指令跟随能力。

AI中文摘要:

当前视觉-语言-动作(VLA)模型的主流方案是将预训练的视觉语言模型(VLM)与单独训练的流匹配动作专家相结合,这使得VLM仅作为上下文编码器而非决策者。本文提出G0.5,一种预训练的自回归VLA模型,其中单个Transformer解码器在单一目标下生成推理与动作令牌。该模型在基础模型规模下可实现的三个关键组件为:可学习的跨 embodiment 动作分词器,能将异构机器人动作映射为共享词汇;原生思维链流,将任务分解、物体定位、动作提示与动作令牌交织;视觉记忆模块,通过视觉编码器注入多秒历史信息。由于推理与动作共享同一组权重,预训练VLM的能力可迁移至物理行为:模型能严格遵循指令,且提示可直接调控动作粒度、任务 horizon 与分布外场景处理,无需额外训练。G0.5在大量机器人数据集与VQA样本上预训练后,在7个独立基准中超越了当前最优模型:R1lite和R1pro机器人的真实世界微调(76.7%,对比π₀.₅的53.3%与GR00T-N1.7的24.4%);2025年BEHAVIOR挑战赛的50项长 horizon 家庭移动操作任务(通用策略下31.4%,对比π₀.₅的26.3%与挑战赛冠军的26.1%);DROID后训练后零样本迁移至未见环境与物体(82.5%);语言跟随的抓取-放置基准LIBERO(98.9%);RoboTwin 2.0(93.3%);SimplerEnv-Bridge(87.3%)

英文摘要:

The prevailing recipe for Vision-Language-Action (VLA) models couples a pretrained VLM with a separately trained flow-matching action expert. This makes the VLM a context encoder rather than a decision-maker. We introduce G0.5, a pretrained autoregressive VLA in which a single transformer decoder emits reasoning and action tokens under a single objective. Three components make this tractable at foundation-model scale: a learnable cross-embodiment action tokenizer that maps heterogeneous robot actions into a shared vocabulary; a native chain-of-thought stream interleaving task decomposition, object grounding, and action hints with action tokens; and a visual memory module that injects multi-second history through the vision encoder. Because reasoning and action share a single set of weights, the pretrained VLM's capabilities carry over to physical behavior: the model follows instructions closely, and prompts directly steer action granularity, task horizon, and out-of-distribution scene handling without further training. Pretrained on a large collection of robot datasets together with VQA samples, G0.5 surpasses state-of-the-art models across 7 independent regimes: real-world fine-tuning on R1lite and R1pro robots (76.7\% vs.\ 53.3\% for $π_{0.5}$ and 24.4\% for GR00T-N1.7), the 2025 BEHAVIOR Challenge on 50 long-horizon household mobile manipulation tasks using a generalist policy (31.4\% vs.\ 26.3\% for $π_{0.5}$ and 26.1\% for the challenge winner), DROID post-training followed by zero-shot transfer to an unseen environment and objects (82.5\%), a language-following Pick-and-Place benchmark, LIBERO (98.9\%), RoboTwin 2.0 (93.3\%), and SimplerEnv-Bridge (87.3\%).

↑