arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.20388cs.ROcs.CV

Navi-Agent:无定位单目导航智能体

Navi-Agent: Unlocalized Monocular Navigation Agent

Wenyuan Xie, Mengyang Hong, Yongzhong Wang, Yanbiao Ji, Yijin Zhou, Shaokai Wu, Shalayiding Sirejiding, Huayi Zhou, Yi-Chao Chen, Ma Ling, Yue Ding, Hongtao Lu

首次发表
浏览论文内容

中文总结 AI 辅助

Navi-Agent提出一种无坐标空间状态(导航拓扑)的零样本VLN-CE智能体,通过视觉观察与运动历史实现自定位、进度验证与恢复,在基准和真实机器人上达到几何约束方法最优性能。

中文摘要 AI 辅助

连续环境中的视觉语言导航(VLN-CE)要求具身智能体在未知环境中执行长时程指令。现有的零样本VLN-CE系统通常通过几何定位或基于坐标的表示来维持空间状态。最近的几何约束导航方法去除了深度和全局一致坐标,但维持持久空间感知以进行位置确认、进度验证和恢复仍然具有挑战性。我们提出Navi-Agent,一种零样本VLN-CE智能体,它从视觉观察和执行的运动历史中构建无坐标空间状态。Navi-Agent将该状态组织为导航拓扑,其中节点表示视觉位置,边表示运动转换。这种表示支持基于观察的近似自定位、任务进度验证和基于视觉重访的恢复。Navi-Agent通过将指令分解为子目标、执行局部视觉导航并通过构建的空间状态验证已访问位置来进行闭环导航。在零样本VLN-CE基准测试和真实世界机器人平台上的实验表明,Navi-Agent在几何约束方法中达到了最先进的性能,同时与依赖几何定位的方法相比仍具有竞争力。

英文摘要

Vision-Language Navigation in Continuous Environments (VLN-CE) requires an embodied agent to execute long-horizon instructions in unknown environments. Existing zero-shot VLN-CE systems typically maintain spatial states through geometric localization or coordinate-based representations. Recent geometry-constrained navigation removes depth and globally consistent coordinates, but maintaining persistent spatial awareness for place confirmation, progress verification, and recovery remains challenging. We present Navi-Agent, a zero-shot VLN-CE agent that constructs a coordinate-free spatial state from visual observations and executed motion histories. Navi-Agent organizes this state as a navigation topology, where nodes represent visual places and edges represent motion transitions. This representation enables observation-based approximate self-localization, task progress verification, and visual revisitation-based recovery. Navi-Agent performs closed-loop navigation by decomposing instructions into sub-goals, executing local visual navigation, and verifying visited places through the constructed spatial state. Experiments on zero-shot VLN-CE benchmark and real-world robot platforms show that Navi-Agent achieves state-of-the-art performance among geometry-constrained methods while remaining competitive with approaches relying on geometric localization.

发表机构

  • Shanghai Jiao Tong University(上海交通大学)
  • Xinjiang University of Finance and Economics(新疆财经大学)
  • Shenzhen University(深圳大学)
  • Southern University of Science and Technology(南方科技大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑