G0.5:用于机器人推理与动作的单自回归流
G0.5: One Autoregressive Stream for Robot Reasoning and Action
- Galaxea(星系科技(Galaxea))
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文提出G0.5,一种单自回归流的VLA模型,通过三个关键组件实现推理与动作统一,在7项基准中超越现有最优模型,展现出良好的泛化与指令跟随能力。
AI中文摘要:
当前视觉-语言-动作(VLA)模型的主流方案是将预训练的视觉语言模型(VLM)与单独训练的流匹配动作专家相结合,这使得VLM仅作为上下文编码器而非决策者。本文提出G0.5,一种预训练的自回归VLA模型,其中单个Transformer解码器在单一目标下生成推理与动作令牌。该模型在基础模型规模下可实现的三个关键组件为:可学习的跨 embodiment 动作分词器,能将异构机器人动作映射为共享词汇;原生思维链流,将任务分解、物体定位、动作提示与动作令牌交织;视觉记忆模块,通过视觉编码器注入多秒历史信息。由于推理与动作共享同一组权重,预训练VLM的能力可迁移至物理行为:模型能严格遵循指令,且提示可直接调控动作粒度、任务 horizon 与分布外场景处理,无需额外训练。G0.5在大量机器人数据集与VQA样本上预训练后,在7个独立基准中超越了当前最优模型:R1lite和R1pro机器人的真实世界微调(76.7%,对比π₀.₅的53.3%与GR00T-N1.7的24.4%);2025年BEHAVIOR挑战赛的50项长 horizon 家庭移动操作任务(通用策略下31.4%,对比π₀.₅的26.3%与挑战赛冠军的26.1%);DROID后训练后零样本迁移至未见环境与物体(82.5%);语言跟随的抓取-放置基准LIBERO(98.9%);RoboTwin 2.0(93.3%);SimplerEnv-Bridge(87.3%)
英文摘要:
The prevailing recipe for Vision-Language-Action (VLA) models couples a pretrained VLM with a separately trained flow-matching action expert. This makes the VLM a context encoder rather than a decision-maker. We introduce G0.5, a pretrained autoregressive VLA in which a single transformer decoder emits reasoning and action tokens under a single objective. Three components make this tractable at foundation-model scale: a learnable cross-embodiment action tokenizer that maps heterogeneous robot actions into a shared vocabulary; a native chain-of-thought stream interleaving task decomposition, object grounding, and action hints with action tokens; and a visual memory module that injects multi-second history through the vision encoder. Because reasoning and action share a single set of weights, the pretrained VLM's capabilities carry over to physical behavior: the model follows instructions closely, and prompts directly steer action granularity, task horizon, and out-of-distribution scene handling without further training. Pretrained on a large collection of robot datasets together with VQA samples, G0.5 surpasses state-of-the-art models across 7 independent regimes: real-world fine-tuning on R1lite and R1pro robots (76.7\% vs.\ 53.3\% for $π_{0.5}$ and 24.4\% for GR00T-N1.7), the 2025 BEHAVIOR Challenge on 50 long-horizon household mobile manipulation tasks using a generalist policy (31.4\% vs.\ 26.3\% for $π_{0.5}$ and 26.1\% for the challenge winner), DROID post-training followed by zero-shot transfer to an unseen environment and objects (82.5\%), a language-following Pick-and-Place benchmark, LIBERO (98.9\%), RoboTwin 2.0 (93.3\%), and SimplerEnv-Bridge (87.3\%).