发表机构
East China Normal University; King Abdullah University of Science and Technology; NeoteAI(华东师范大学; 阿卜杜拉国王科技大学; NeoteAI)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对VLN中语义推理与空间执行脱节的问题,提出GroundingVLN,以视觉锚定为共享接口,结合锚定推理、像素目标预测及GEAR强化学习,在极低数据下达到SOTA性能。
AI 中文摘要
尽管视觉语言模型(VLM)具备强大的视觉理解和推理能力,现有的视觉语言导航(VLN)智能体却难以将语义推理与空间执行相连接。这种连接中存在两个相互耦合的缺口:中间推理过程没有明确锚定到视觉证据上,且高层决策缺乏精确的空间目标来指导低层运动。认知科学表明,人类导航通过将认知锚定到相关地标、并引导运动朝向空间目标,以层级方式弥合这些层次。受此原理启发,我们提出GroundingVLN,将视觉锚定作为推理与行动之间的共享接口。GroundingVLN首先通过锚定进行推理,在结构化推理过程中将任务相关的视觉证据锚定到精确的图像位置;随后通过锚定进行行动,预测一个与进度对齐的像素目标,由几何规划器将其转化为基本动作。为学习这些能力,我们构建了GroundingCOTVLN-188K数据集,包含时间对齐的锚定推理轨迹,并引入锚定与执行感知强化学习(GEAR),将锚定推理和空间决策与下游执行对齐。实验表明,GroundingVLN在R2R-CE上达到69.9%的SR,在RxR-CE上达到75.1%的SR,实现了最先进的性能,且样本效率极高,仅使用最强基线0.9%的训练数据。它还能跨数据集强泛化,仅用R2R训练即在RxR-CE上达到59.9%的SR,比最强基线高出20.1%。
英文摘要
Although vision-language models (VLMs) possess strong visual understanding and reasoning capabilities, existing vision-and-language navigation (VLN) agents struggle to connect semantic reasoning with spatial execution. Two coupled gaps remain in this connection, as intermediate reasoning is not explicitly anchored to visual evidence and high-level decisions lack precise spatial goals to guide low-level motion. Cognitive science suggests that human navigation bridges these levels hierarchically by anchoring cognition to relevant landmarks and guiding locomotion toward spatial goals. Motivated by this principle, we propose GroundingVLN, which uses visual grounding as a shared interface between reasoning and action. GroundingVLN first reasons with grounding by anchoring task-relevant visual evidence to precise image locations throughout structured reasoning. It then acts through grounding by predicting a progress-aligned pixel goal that a geometric planner translates into primitive actions. To learn these capabilities, we construct GroundingCOTVLN-188K, a dataset of temporally aligned grounded reasoning traces, and introduce Grounded and Execution-Aware Reinforcement Learning (GEAR), which aligns grounded reasoning and spatial decisions with downstream execution. Experiments demonstrate that GroundingVLN achieves state-of-the-art performance (69.9% SR on R2R-CE and 75.1% SR on RxR-CE) with high sample efficiency, using just 0.9% as much training data as the strongest baseline. It also generalizes strongly across datasets, attaining 59.9% SR on RxR-CE when trained solely on R2R, a gain of 20.1% over the strongest baseline. Code and models will be released after review. Code is available at [https://github.com/Teacher-Tom/GroundingVLN](https://github.com/Teacher-Tom/GroundingVLN).