面向开放词汇实例导航的基于对象-路径图的拓扑表示
A Topological Representation with Object-Path Graphs for Open-Vocabulary Instance Navigation
- The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
提出对象-路径图统一开放词汇语义推理与拓扑导航,实现无需稠密度量地图的全局规划与局部执行,在HM3D和Replica上验证有效。
AI中文摘要:
视觉语言导航要求具身智能体利用自然语言指令和视觉观察在环境中导航。现有方法通常将导航分解为顺序的语言引导决策,或依赖缺乏先验环境知识的在线探索。场景图表示提供了紧凑的语义记忆,但与下游导航仍然脱节,后者仍依赖于稠密度量地图。为弥合这一差距,我们提出了一种对象-路径图,将开放词汇语义推理与拓扑导航统一起来。所提出的表示在单一轻量级拓扑框架内联合支持语义定位、基于图的定位和导航。基于该图,我们引入了一种导航策略,该策略通过轻量级节点定位和语义视觉伺服,将全局路径规划与局部节点间执行相结合,从而无需稠密度量重建即可直接在图上进行导航。在HM3D和Replica上的实验表明,通过所提出的分层图结构,在开放词汇对象定位方面具有竞争力的性能,同时实现了有效的导航性能。真实机器人实验进一步验证了所提出框架的实用性。
英文摘要:
Vision-language navigation requires embodied agents to navigate environments using natural language instructions and visual observations. Existing approaches typically decompose navigation into sequential language-guided decisions or rely on online exploration without prior environmental knowledge. Scene graph representations offer compact semantic memory but remain decoupled from downstream navigation, which still depends on dense metric maps. To close this gap, we propose an object--path graph that unifies open-vocabulary semantic reasoning with topological navigation. The proposed representation jointly supports semantic grounding, graph-based localization, and navigation within a single lightweight topological framework. Building on this graph, we introduce a navigation strategy that combines global path planning with local inter-node execution through lightweight node localization and semantic visual servoing, enabling navigation directly over the graph without dense metric reconstruction. Experiments on HM3D and Replica demonstrate competitive performance in open-vocabulary object grounding through the proposed hierarchical graph structure, while achieving effective navigation performance. Real-world robot experiments further validate the practicality of the proposed framework.