arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

自动驾驶中视觉-语言-动作的显式几何思维链

Explicit Geometric Chain-of-Thought for Vision-Language-Action in Autonomous Driving

Xingtai Gui, Yucheng Zhou, Dongqian Guo, Jiahao Gong, Feiyang Tan, Jianbing Shen

arXiv 2610.10390首次发表:更新:

发表机构

University of Macau; Afari Intelligent Drive(澳门大学; Afari智能驾驶)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对VLA模型在自动驾驶中2D推理与3D动作不匹配的问题,提出GeoCoTDrive显式几何思维链框架,通过2D grounding检索3D先验并交织入自回归上下文,显著提升安全关键规划性能。

AI 中文摘要

视觉-语言-动作(VLA)模型已成为自动驾驶领域一种前景广阔的研究范式。然而,现有的VLA模型仍存在一个根本性的不匹配:驾驶动作需要精确的3D几何线索,而视觉-语言理解与推理则主要在2D语义空间中进行。在本文中,我们提出了GeoCoTDrive,一个显式的几何思维链框架,它以规划为导向的方式对几何信息进行 grounding。GeoCoTDrive遵循“先用2D思考,再用专用3D先验驱动”的范式。它首先对与决策关键线索对应的2D区域进行 grounding,然后通过在已 grounding 区域内从几何基础模型中采样特征来检索局部3D先验。这些局部几何特征被交织到自回归上下文中,以支持轨迹生成。为了监督这一过程,我们引入了规划相关 grounding,这是一种新的区域级 grounding 任务,专注于直接影响自我规划决策的局部空间线索,并构建了PlanningGrounding数据集,以赋予VLA模型规划导向的 grounding 能力。在多个端到端自动驾驶基准上的实验表明,GeoCoTDrive持续提升了安全关键规划性能,证明了显式几何思维链过程对基于VLA的规划的有效性。

英文摘要

Vision-language-action~(VLA) models have emerged as a promising paradigm for autonomous driving. However, existing VLA models still suffer from a fundamental mismatch: driving actions require precise 3D geometric cues, while visual-language understanding and reasoning are largely conducted in a 2D semantic space. In this paper, we propose GeoCoTDrive, an explicit geometric chain-of-thought framework that grounds geometry in a planning-oriented manner. GeoCoTDrive follows a think with 2D first, drive with dedicated 3D priors paradigm. It first grounds 2D regions corresponding to decision-critical cues, and then retrieves localized 3D priors by sampling features from a geometric foundation model within the grounded regions. These localized geometric features are interleaved into the autoregressive context to support the trajectory generation. To supervise this process, we introduce planning-relevant grounding, a new region-level grounding task that focuses on local spatial cues directly affecting ego planning decisions, and construct the PlanningGrounding dataset to endow VLAs with planning-oriented grounding capability. Experiments across multiple end-to-end autonomous driving benchmarks show that GeoCoTDrive consistently improves safety-critical planning performance, demonstrating the effectiveness of the explicit geometric chain-of-thought process for VLA-based planning.

Comments21 pages, 9 figures. The code is available at https://github.com/TabGuigui/GeoCoTDrive

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑