AI 中文总结
该研究针对VLA模型在复杂驾驶场景中规划性能不足的问题,提出可即插即用的Geo-VLA框架,结合自研Geo-QA数据集,在NAVSIM v1上实现单目VLA规划器的新SOTA。
AI 中文摘要
视觉-语言-动作(VLA)模型已通过利用基础模型进行语义推理和长尾泛化,在端到端自动驾驶领域取得了进展。然而,由于仅基于图像的表示无法充分捕捉规划所需的道路几何与拓扑信息,它们在复杂驾驶环境中的规划性能仍然有限。本文提出Geo-VLA,这是一种可即插即用的框架,通过学习几何感知的视觉表示来增强VLA模型。训练期间,Geo-VLA内化几何地图语义以强化道路结构表示,推理时无需高精地图(HD maps)或额外车道信息。为支撑该方法,我们引入Geo-QA,一个聚焦几何的问答数据集,通过对比学习与指令调优将道路几何注入视觉-语言表示。在NAVSIM v1上的实验表明,Geo-VLA可持续改进具有不同动作生成架构的VLA规划器,达到92.1的PDMS指标,成为单目VLA规划器中的新 State-of-the-Art(SOTA)。
英文摘要
Vision-language-action (VLA) models have advanced end-to-end autonomous driving by leveraging foundation models for semantic reasoning and long-tail generalization. However, their planning performance remains limited in complex driving environments because image-only representations inadequately capture planning-relevant road geometry and topology. In this paper, we propose Geo-VLA, a plug-and-play framework that enhances VLA models by learning geometry-aware visual representations. During training, Geo-VLA internalizes geometric map semantics to strengthen road-structure representations, while requiring no HD maps or additional lane information during inference. To support this approach, we introduce Geo-QA, a geometry-focused question-answering dataset that injects road geometry into vision-language representations through contrastive learning and instruction tuning. Experiments on NAVSIM v1 demonstrate that Geo-VLA consistently improves VLA planners with distinct action-generation architectures, achieving 92.1 PDMS and establishing a new state-of-the-art among single-camera VLA planners.