arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.39915cs.CVcs.RO

NavHarness:智能体视觉语言导航的自适应目标

NavHarness: Adaptive Goals for Agentic Vision-Language Navigation

  • Harbin Institute of Technology (Shenzhen)(哈尔滨工业大学(深圳))
  • Pengcheng Laboratory(鹏城实验室)

机构由 AI 辅助整理,请以论文原文为准。

Haoxiang Shi, Zaijing Li, Muhe Ding, Xiang Deng, Yaowei Wang, Liqiang Nie

AI总结:

NavHarness提出一种智能体VLN框架,通过目标、验证、记忆和视觉运动智能体协同,实现自适应目标设定与历史压缩,在R2R-CE和RxR-CE上取得83.3%成功率和1.51米导航误差。

AI中文摘要:

视觉语言导航(VLN)要求具身智能体根据指令和观察生成动作。通用多模态智能体为此任务提供了有前景的基础,但选择合理的局部动作并不能确保执行与预期路线保持一致,尤其是在长时程任务中。此外,累积的交互历史增加了后续决策所需的输入,导致显著的推理开销。为此,我们引入了 \nmethod,一个智能体VLN框架,包括一个目标智能体(Goal Agent)为局部动作设定自适应目标,一个验证智能体(Verify Agent)动态验证目标是否完成,一个记忆智能体(Memory Agent)用于多模态上下文压缩,以及一个视觉运动智能体(Visuomotor Agent)执行自适应目标。具体而言,目标智能体根据指令、当前观察和执行历史制定自适应目标。然后视觉运动智能体执行导航动作以实现每个目标,而验证智能体使用特定于目标的验证问题动态评估观察到的结果是否满足预期的完成条件。验证完成的目标随后标记边界,使记忆智能体压缩相应的多模态交互历史,同时保留后续导航所需的信息。我们在R2R-CE和RxR-CE上评估导航,检查三种模型骨干的框架变体,并研究执行过程中的上下文演变。对于真实世界评估,\nmethod在八条具有挑战性的路线上(每条评估三次)实现了83.3%的成功率和1.51米的导航误差。

英文摘要:

Vision-Language Navigation (VLN) requires embodied agents to generate actions based on instructions and observations. General-purpose multimodal agents offer a promising basis for this task, but selecting plausible local actions does not ensure that execution remains consistent with the intended route, particularly in long-horizon tasks. Moreover, the accumulated interaction history increases the input required for subsequent decisions, resulting in a significant inference overhead. To this end, we introduce \method, an Agentic VLN framework that includes a Goal Agent that sets adaptive goals for local actions, a Verify Agent that dynamically verifies whether a goal has been completed, a Memory Agent for multimodal context compression, and a Visuomotor Agent to execute adaptive goals. Specifically, the Goal Agent formulates adaptive goals based on the instruction, current observation, and execution history. Then the Visuomotor Agent executes navigation actions to achieve each goal, while the Verify Agent uses a goal-specific verification question to dynamically assess whether the observed outcomes satisfy the intended completion condition. Verified goal completion then marks a boundary for the Memory Agent to compress the corresponding multimodal interaction history while preserving information needed for subsequent navigation. We evaluate navigation on R2R-CE and RxR-CE, examine framework variants across three model backbones, and study context evolution during execution. For Real-World evaluation, \method achieves 83.3\% success and 1.51\,m navigation error across eight challenging routes evaluated three times each.

补充信息

↑