arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.12630cs.ROcs.CV

用于视觉语言导航的实例增强语义地图

Instance-Enriched Semantic Maps for Visual Language Navigation

Jiho Hong, Eunae Kang, Sanghyun Kim, Young-Sik Shin

首次发表
浏览论文内容

中文总结 AI 辅助

研究视觉语言导航问题,提出实例增强语义地图框架,通过实例级二维半丰富信息映射、基于大语言模型的鲁棒查询处理和存储高效的语义表示,提升导航性能,在相关实验中取得优于基线的结果。

中文摘要 AI 辅助

视觉语言导航旨在让具身智能体根据自然语言指令在复杂环境中导航。现有方法构建语义空间地图并利用大语言模型进行推理决策,但缺乏实例级对象细节和对不同用户查询的鲁棒性。为此提出实例增强语义地图框架,包括通过全景分割构建实例级二维半丰富信息地图、基于大语言模型的鲁棒查询处理及存储高效的语义表示。实验表明该方法在预测归一化曲线下面积等指标上优于三维基线,导航实验中对象检索和导航成功率也有显著提升。

英文摘要

Visual Language Navigation (VLN) aims to enable an embodied agent to navigate complex environments by following natural language instructions. Recent approaches build semantic spatial maps and leverage Large Language Models (LLMs) for reasoning and decision making. Despite these advances, existing systems lack instance-level object detail and robustness to diverse user queries, limiting reliable navigation in complex indoor environments. To address these limitations, we propose Instance-Enriched Semantic Maps, a unified framework with three key contributions: (1) Instance-level two-and-a-half-dimensional (2.5D) rich information mapping that constructs maps from color and depth observations via open-vocabulary panoptic segmentation, preserving vertical distinctions and capturing small objects, while storing diverse semantic attributes and natural language captions enriched with room-level context. (2) Robust query processing via LLM-based target selection, which dynamically routes queries across type-specialized experts and integrates their outputs through score-level fusion, enabling consistent goal selection across diverse query formulations. (3) Storage-efficient semantic representation that achieves approximately 96% reduction compared to three-dimensional (3D) scene-graph approaches while preserving sufficient spatial information for navigation. The proposed 2.5D representation outperforms the 3D baseline by over 27% in prediction-normalized Area Under the Curve (AUC). In navigation experiments, our method achieves over 17% improvement in object retrieval and over 23% in navigation success compared to the baseline across diverse query types. The project page is available at https://rcilab.github.io/iesm_vln.

发表机构

  • Department of Mechanical Engineering, Kyung Hee University(庆熙大学机械工程系)
  • Advanced Institutes of Convergence Technology (AICT)(融合技术高级研究院)
  • School of Mechanical Engineering, Kyungpook National University(庆北国立大学机械工程学院)

机构由 AI 辅助整理,请以论文原文为准。

↑