arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.12285cs.ROcs.CV

AnchorVLN:面向开放词汇导航的几何锚定视觉-语言基础推理

AnchorVLN: Geometry-Anchored Vision-Language Grounding Reasoning for Open-Vocabulary Navigation

Long Giang Vu, Chengkai Yao, Yuxin Liu, FNU Aryan, Rajath Chandrashekar Aralikatti

首次发表
浏览论文内容

中文总结 AI 辅助

AnchorVLN提出语义由VLM提出、几何决定度量的规则,通过MCP工具分离语义与几何,在CMU VLN挑战2026上指令遵循达64.4%,物体引用中位误差降至2.48米。

中文摘要 AI 辅助

在未知室内环境中的视觉-语言导航(VLN)在现实世界机器人技术中具有实用价值,其中智能体必须遵循自然语言指令、定位物体并回答空间问题,而无需预先构建地图或固定物体词汇表。多模态视觉-语言模型(VLM)提供了强大的开放词汇基础能力和零样本推理能力,但难以直接从图像中输出可靠的度量量,如距离、方位和比较性空间关系。现有方法通过将几何信息融入手工设计的流程或要求模型输出航点来解决这一问题,这需要针对不同机器人、任务或词汇表修改控制栈。我们提出AnchorVLN,一个基于简单规则的开放词汇VLN系统:VLM提出语义;几何决定度量。该系统实现为EMBODIED-NAV-MCP,一个由VLM智能体通过一组紧凑的可调用工具驱动的模型上下文协议(MCP)服务器。由于没有工具接受以米为单位的距离或以弧度为单位的方位,该模式在无需修改下游自主栈的情况下强制执行语义-几何边界。我们对CMU视觉-语言导航挑战赛2026的两项任务进行了基准测试:15个场景中的30个指令遵循问题和一组冻结的45个物体引用问题。完整系统在指令遵循上达到64.4%,在没有控制器建模的情况下下降13.3个百分点(t=2.77)。在物体引用上,几何锚定在45个问题中的10个上超过了挑战重叠阈值,而直接坐标估计为45个中的0个,将中位中心误差从3.37米降低到2.48米。

英文摘要

Vision-Language Navigation (VLN) in unseen indoor environments is useful in real-world robotics, where an agent must follow natural-language instructions, locate objects, and answer spatial questions without a pre-built map or fixed object vocabulary. Multimodal vision-language models (VLMs) provide strong open-vocabulary grounding and zero-shot reasoning, but struggle to emit reliable metric quantities such as range, bearing, and comparative spatial relations directly from images. Existing approaches address this by folding geometry into hand-engineered pipelines or asking models to output waypoints, requiring changes to the control stack for different robots, tasks, or vocabularies. We introduce AnchorVLN, an open-vocabulary VLN system built on a simple rule: the VLM proposes semantics; geometry decides metrics. It is realised as EMBODIED-NAV-MCP, a Model Context Protocol (MCP) server driven by a VLM agent through a compact set of callable tools. Since no tool accepts distance in metres or bearing in radians, the schema enforces the semantic-geometry boundary without modifying the downstream autonomy stack. We benchmark both tasks of the CMU Vision-Language Navigation Challenge 2026: 30 instruction-following questions over 15 scenes and a frozen 45-question object-reference set. The full system achieves 64.4 percent on instruction following, dropping by 13.3 percentage points without controller modeling (t = 2.77). On object reference, geometric anchoring clears the challenge overlap threshold on 10 of 45 questions, versus 0 of 45 for direct coordinate estimation, reducing median center error from 3.37 m to 2.48 m.

↑