HAM-VLN:利用分层智能体记忆实现零样本视觉与语言导航
HAM-VLN: Harnessing Hierarchical Agentic Memory for Zero-Shot Vision-and-Language Navigation
浏览论文内容
中文总结 AI 辅助
本研究提出HAM-VLN零样本VLN方法,通过分层智能体记忆缓解长程导航的记忆与推理瓶颈,提升导航指标并缩短上下文长度,在多个基准数据集上取得优异零样本性能。
中文摘要 AI 辅助
视觉与语言导航(VLN)使机器人能够在从未见过的环境中遵循指令导航。近年来,一种无需训练的范式应运而生:机器人通过查询多模态大语言模型(multimodal LLM)来理解自身观测并规划下一步行动。然而,基于图像流或密集地图的长程导航不可避免地会引发不断加剧的记忆与推理瓶颈。我们提出HAM-VLN,这是一种与决策耦合、由智能体生成的记忆,为机器人配备了持久且基于深度的世界图。在用于选择下一步行动的同一模型调用中,HAM-VLN还会记录语义与反思信息,包括房间类型、物体、导航进度及失败记录。近期的路径点会在有限窗口内逐字保留,而较早的历史信息仅会通过相关性、时效性和显著性评分的检索,以及单跳拓扑扩展重新进入上下文。该设计无需在每个路径点决策之外调用额外的大语言模型。与先前方法相比,HAM-VLN不仅提升了多项导航指标,还将上下文长度缩短了65%以上。具体而言,HAM-VLN在VLN-CE R2R上实现61.0%的成功率(SR),在VLN-CE RxR上实现52.7%的SR,在HM3D-v2 ObjectNav上实现79.7%的SR,且全程无需任何训练。
英文摘要
Vision-and-language navigation (VLN) enables robots to follow instructions in previously unseen environments. Recently, a training-free paradigm has emerged: the robot queries a multimodal LLM to understand its observations and plan the next action. However, long-horizon navigation based on either image streams or dense map inevitably introduces a growing memory and reasoning bottleneck. We present HAM-VLN, a decision-coupled, agent-authored memory that equips the robot with a persistent, depth-grounded world graph. In the same model call used to select the next action, HAM-VLN also records semantic and reflective information---including room type, objects, navigation progress, and failure notes. Recent waypoints remain verbatim within a bounded window, while older history re-enters the context only through retrieval scored by relevance, recency, and salience, together with one-hop topological expansion. This design requires no additional LLM calls beyond the per-waypoint decision. Compared to previous methods, HAM-VLN not only improves various navigation metrics but also reduces the context length by more than 65%. Specifically, HAM-VLN achieves 61.0% Success Rate (SR) on VLN-CE R2R, 52.7% SR on VLN-CE RxR, and 79.7% SR on HM3D-v2 ObjectNav without any training.