发表机构
National Engineering Research Center of Electric Vehicles, Beijing Institute of Technology; Shenzhen Automotive Research Institute, Beijing Institute of Technology; Shenzhen Jiguangzhijie Technology Co., Ltd.; Nanyang Technological University(北京理工大学电动车辆国家工程研究中心; 北京理工大学深圳汽车研究院; 深圳极光智界科技有限公司; 南洋理工大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
GRAVA通过GRA框架统一接地、推理与动作生成,利用轨迹锚定图和渐进训练,在NAVSIM上达到最先进性能,显著提升驾驶合规性与闭环得分。
AI 中文摘要
驾驶视觉-语言-动作(VLA)模型在行动前越来越多地进行推理,但其中间推理往往缺乏对物理场景证据的强接地性,且与可执行行为之间的连接松散。我们提出GRAVA,一个围绕接地推理到动作(GRA)构建的框架,该框架在单一自回归流中统一了接地、推理和动作生成。GRA将动作相关的语言引用关联到2D视觉区域和以自我为中心的物理状态,在轨迹锚定的类型化图中组织对象交互和决策,并将此结构序列化为接地推理。一个单一的VLM生成此推理,随后生成紧凑的可执行规划器动作,该动作被确定性解码为连续轨迹。我们进一步引入一个智能体GRA数据构建流程,该流程结合前向场景接地与后向轨迹锚定,并利用它构建GR-NavSim,包含220万个接地问答对和7万条GRA推理轨迹。一种渐进式训练策略通过预训练发展接地认知,通过模仿建立推理到动作的接口,并通过强化学习和探索改进驾驶行为。使用约60%的可用人类驾驶演示进行动作监督,GRAVA-8B在完整NAVSIM基准上实现了纯自回归驾驶模型中的最先进性能。在内部长尾基准上,完整的GRA相比仅动作预测,关键对象合规性和闭环驾驶得分分别提高了19.3%和20.5%。这些结果表明,在从接地推理到可执行动作生成的过程中保留动作相关的物理证据具有益处。
英文摘要
Driving vision-language-action (VLA) models increasingly reason before acting, but their intermediate reasoning is often weakly grounded in physical scene evidence and loosely connected to executable behavior. We present GRAVA, a framework built around Grounded Reasoning-to-Action (GRA), which unifies grounding, reasoning, and action generation in a single autoregressive stream. GRA links action-relevant language references to 2D visual regions and ego-centric physical states, organizes object interactions and decisions in a trajectory-anchored typed graph, and serializes this structure into grounded reasoning. A single VLM generates this reasoning followed by a compact Executable Planner action that is deterministically decoded into a continuous trajectory. We further introduce an agentic GRA data construction pipeline that combines forward scene grounding with backward trajectory anchoring, and use it to build GR-NavSim with 2.2M grounded question-answer pairs and 70K GRA reasoning traces. A progressive training strategy develops grounded cognition through pre-training, establishes the reasoning-to-action interface through imitation, and improves driving behavior through reinforcement learning and exploration. Using about 60% of the available human driving demonstrations for action supervision, GRAVA-8B achieves state-of-the-art performance among purely autoregressive driving models on the full NAVSIM benchmark. On an internal long-tail benchmark, full GRA improves key-object compliance and Closed-loop Driving Score by 19.3% and 20.5% over action-only prediction, respectively. These results show the benefit of preserving action-relevant physical evidence from grounded reasoning through executable action generation.
Comments23 pages. Code: https://github.com/AhernResearch/grava