arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

VTM-Nav:用于跨情节对象目标导航的分层视觉拓扑记忆

VTM-Nav: Harnessing Cross-Episode Experience for Object-Goal Navigation with Hierarchical Visual-Topological Memory

Xiaoran Xu, Yupeng Wu, Tianyu Xue, Yifan Xu, Xuanran Dong, Xiaoshan Yang, Changsheng Xu

arXiv 2607.14514首次发表:更新:

发表机构

MAIS, Institute of Automation, Chinese Academy of Sciences; School of Advanced Interdisciplinary Sciences, University of Chinese Academy of Sciences; Tsinghua University; University of Chinese Academy of Sciences(中国科学院自动化研究所模式识别国家重点实验室; 中国科学院大学交叉科学学院; 清华大学; 中国科学院大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究跨情节对象目标导航问题,提出无训练的VLM导航框架\method,其带有分层视觉拓扑记忆,通过粗到细匹配检索经验,在三个基准测试中性能最佳,证明跨数据集结构化视觉拓扑经验复用有效且鲁棒。

AI 中文摘要

对象目标导航要求实体智能体在室内环境中定位并到达指定对象类别的一个实例。近期无训练方法利用视觉语言模型进行开放词汇语义推理,但通常在情节协议下评估,即每集后重置所有场景特定状态。我们引入跨情节对象目标导航,智能体在同一场景中反复操作,仅保留自身获取的经验且模型参数固定。为支持经验复用,我们提出一种无训练的VLM导航框架\method,带有持久分层视觉拓扑记忆(VTM)。VTM在房间和对象级别组织场景知识,通过粗到细匹配检索相关经验,仅在与当前观察一致时提供记忆作为软指导。保守执行防护进一步减轻振荡、受阻运动和过早停止。在受控的同一场景协议下,我们在三个基准测试HM3D v0.1、HM3D v0.2和MP3D上评估\method,并与增强了跨情节文本记忆的强化WMNav基线进行比较,同时保持VLM主干和动作管道相同。\method在所有三个基准测试中取得最佳性能,证明了跨数据集结构化视觉拓扑经验复用的有效性和鲁棒性。

英文摘要

Training-free ObjectNav agents increasingly use vision-language models (VLMs), yet typically discard acquired scene knowledge after each request. We study cross-episode ObjectNav, where each request is an independently initialized, single-goal episode and only self-acquired, scene-scoped memory persists across episodes. We ask whether an agent with fixed model parameters and navigation components can reuse such experience without retraining or oracle information. We introduce \method, a training-free framework with a persistent Hierarchical Visual-Topological Memory (VTM). VTM uses a coarse room topology to index room-owned visual memories, distinguishes in-room from remote-visible evidence, and retains successful approach cues. For each request, VTM-Nav re-localizes the agent in accumulated scene structure, retrieves target-relevant records from plausible rooms, and grounds memory guidance in candidates derived from the current observation. A conservative execution guard further handles local failures. Under matched 40-step comparisons, VTM-Nav exceeds the memory-reset WMNav control by 4.6, 2.0, and 0.8 SR points on HM3D v0.1, HM3D v0.2, and MP3D, respectively, with comparable or higher SPL. On HM3D, it also exceeds WMNav harnessed by textual memory by 3.1 and 5.5 SR points. These results demonstrate effective reuse of cross-episode scene experience through hierarchical visual-topological memory.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑