arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.22226cs.RO

用于户外环境的基于几何目标定位的离线视觉语言导航

Offline Vision-Language Navigation with Geometric Goal Localization for Outdoor Environments

Ali Salmasi, Xianjia Yu, Tomi Westerlund

首次发表
浏览论文内容

中文总结 AI 辅助

研究针对户外环境下的离线视觉语言导航,通过对17个可边缘部署的SLM与4个在线API进行基准测试,提出轻量级混合语义几何目标定位框架并集成到Edge-BehAV中,提升了性能,降低目标距离误差,实现无网络连接下的有效导航。

中文摘要 AI 辅助

基于基础模型的视觉语言导航(VLN)推动了自主机器人导航发展。然而,现有VLN系统严重依赖云托管基础模型,限制了其在网络不可用及需要可靠度量目标定位场景的应用。本文有三个贡献:一是对17个可边缘部署的小语言模型(SLM)与4个在线应用程序编程接口(API)进行系统基准测试;二是提出轻量级混合语义几何目标定位框架;三是将这些进展集成到Edge-BehAV中。实验结果表明,最佳离线SLM在无网络连接时运行速度快约9倍且指令分解性能与最强云API相当,目标定位框架降低了平均目标距离误差,完整系统在32次闭环户外试验中有31次成功。

英文摘要

Foundation-model-based vision-language navigation (VLN) has advanced autonomous robot navigation by enabling robots to interpret natural-language instructions, identify semantic goals, and follow user-specified behavioral rules. However, existing VLN systems rely heavily on cloud-hosted foundation models for language understanding and semantic grounding, limiting their applicability where network connectivity is unavailable and reliable metric goal localization is required. Although recent small language models (SLMs) enable fully onboard inference, their suitability for navigation instruction decomposition has not been systematically evaluated. This paper makes three contributions toward fully onboard VLN for outdoor environments. First, we present the first systematic benchmark of 17 edge-deployable SLMs against 4 online APIs for robotic navigation instruction decomposition, evaluating accuracy and latency on human-annotated instructions across three computing platforms and providing practical guidance for selecting onboard language models. Second, we propose a lightweight hybrid semantic-geometric goal localization framework that combines open-vocabulary object detection, prompted segmentation, and LiDAR geometry to estimate metric goals, while maintaining visual bearing guidance when reliable geometric observations are unavailable. Third, we integrate these advances into Edge-BehAV, a fully onboard extension of the BehAV architecture that enables cloud-independent behavior-guided navigation. Experimental results show that the best offline SLM matches the instruction decomposition performance of the strongest cloud API while running approximately 9x faster and without network connectivity. The proposed goal localization framework reduces mean goal-distance error from 2.05 m to 0.20 m at lower computational cost, and the complete system succeeds in 31 of 32 closed-loop outdoor trials.

↑