arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2603.17670cs.RO

AgentVLN:迈向具有代理能力的视觉-语言导航

AgentVLN: Towards Agentic Vision-and-Language Navigation

  • Nanjing University of Aeronautics and Astronautics(南京航空航天大学)
  • Shandong University(山东大学)
  • Zhejiang University(浙江大学)

机构由 AI 辅助整理,请以论文原文为准。

Zihao Xin, Wentong Li, Yixuan Jiang, Ziyuan Huang, Bin Wang, Piji Li, Jianke Zhu, Jie Qin, Shengjun Huang

更新

AI总结:

本文提出AgentVLN框架,通过部分可观测半马尔可夫决策过程建模视觉-语言导航,结合VLM-as-Brain范式和跨空间表示映射,解决多级表示不一致问题,实现高效轻量级导航部署。

AI中文摘要:

视觉-语言导航(VLN)要求一个具身体验的代理将复杂的自然语言指令转化为在未知环境中长距离导航的路径。尽管视觉-语言模型(VLMs)在2D语义理解方面表现优异,但当前的VLN系统仍受限于有限的空间感知、2D-3D表示不匹配以及单目尺度模糊。本文提出AgentVLN,一种新颖且高效的具身体验导航框架,可部署在边缘计算平台上。我们将VLN建模为部分可观测半马尔可夫决策过程(POSMDP),并引入VLM-as-Brain范式,通过插件式技能库将高层语义推理与感知和规划解耦。为解决多级表示不一致问题,我们设计了跨空间表示映射,将感知层的3D拓扑路径点投影到图像平面,产生像素对齐的视觉提示。在此基础上,我们整合了上下文感知的自我纠正和主动探索策略,以恢复遮挡并抑制长轨迹中的误差累积。为进一步解决无结构环境中指令的空间模糊性,我们提出了查询驱动的感知思维链(QD-PCoT)方案,使代理具备元认知能力,主动寻求几何深度信息。最后,我们构建了AgentVLN-Instruct,一个大规模指令微调数据集,包含动态阶段路由,条件于目标可见性。大量实验表明,AgentVLN在长距离VLN基准上一致优于先前最先进的方法(SOTA),提供了一种实用的轻量级部署下一代具身体验导航模型的范式。代码:https://github.com/Allenxinn/AgentVLN。

英文摘要:

Vision-and-Language Navigation (VLN) requires an embodied agent to ground complex natural-language instructions into long-horizon navigation in unseen environments. While Vision-Language Models (VLMs) offer strong 2D semantic understanding, current VLN systems remain constrained by limited spatial perception, 2D-3D representation mismatch, and monocular scale ambiguity. In this paper, we propose AgentVLN, a novel and efficient embodied navigation framework that can be deployed on edge computing platforms. We formulate VLN as a Partially Observable Semi-Markov Decision Process (POSMDP) and introduce a VLM-as-Brain paradigm that decouples high-level semantic reasoning from perception and planning via a plug-and-play skill library. To resolve multi-level representation inconsistency, we design a cross-space representation mapping that projects perception-layer 3D topological waypoints into the image plane, yielding pixel-aligned visual prompts for the VLM. Building on this bridge, we integrate a context-aware self-correction and active exploration strategy to recover from occlusions and suppress error accumulation over long trajectories. To further address the spatial ambiguity of instructions in unstructured environments, we propose a Query-Driven Perceptual Chain-of-Thought (QD-PCoT) scheme, enabling the agent with the metacognitive ability to actively seek geometric depth information. Finally, we construct AgentVLN-Instruct, a large-scale instruction-tuning dataset with dynamic stage routing conditioned on target visibility. Extensive experiments show that AgentVLN consistently outperforms prior state-of-the-art methods (SOTA) on long-horizon VLN benchmarks, offering a practical paradigm for lightweight deployment of next-generation embodied navigation models. Code: https://github.com/Allenxinn/AgentVLN.

补充信息

↑