发表机构
University College London; University of London; China Merchants Group; LionRock AI Lab; Shenzhen University(伦敦大学学院; 伦敦大学; 招商局集团; 狮子岩人工智能实验室; 深圳大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
MarvisNav通过将探索记忆直接投影到视觉路径选择上,使VLM联合评估目标相关性和探索状态,无需策略训练,在HM3D上取得最先进性能,并显著减少VLM调用。
AI 中文摘要
当寻找一个物体时,人们会同时考虑可能的目标位置和已经探索过的地方来做出下一步行动。当前视图可以提示与地点相关的记忆,将目标相关性和先前探索置于同一空间上下文中。然而,在许多零样本目标导航(ZSON)方法中,视觉语言模型(VLMs)从自我中心图像推断有希望的搜索区域,而探索历史则被单独表示,例如作为文本或地图。这种分离要么需要额外的融合步骤,要么使记忆与路径选择之间的对应关系隐式化,留给VLM去恢复。相反,我们让探索记忆直接可见于视觉路径选择上。我们提出MarvisNav,一个ZSON框架,它维护一个拓扑图,并将候选节点及其探索状态投影到自我中心视图上,作为携带记忆的视觉路径选择。这些状态捕获了超越二元访问的局部探索进度。通过将探索状态直接绑定到每个视觉候选,MarvisNav使VLM能够联合评估目标相关性和探索状态,而无需单独的事后融合或重排序阶段。无需策略训练,MarvisNav在HM3D上达到最先进性能(81.2% SR和42.5% SPL),同时在MP3D上保持竞争力。它还以远少于代表性VLM方法的VLM调用次数(例如,WMNav的7.5%)超越它们。跨多样场景的真实机器人实验进一步验证了其实用可部署性。超越MarvisNav,我们的研究表明记忆表示塑造VLM决策和ZSON性能,强调有效记忆使用不仅取决于其可用性,还取决于其表示方式。代码和项目页面将在 \u200b\u200bhttps:// 提供。
英文摘要
When searching for an object, people choose their next move by considering both likely target locations and places already explored. The current view can cue place-associated memories, bringing target relevance and prior exploration into the same spatial context. In many zero-shot object navigation (ZSON) methods, however, vision-language models (VLMs) infer promising search areas from egocentric images, while exploration history is represented separately, e.g., as text or maps. This separation either requires an additional fusion step or leaves the correspondence between memory and route choices implicit for the VLM to recover. We instead make exploration memory directly visible on visual route choices. We propose MarvisNav, a ZSON framework that maintains a topological graph and projects candidate nodes together with their exploration states onto egocentric views as memory-bearing visual route choices. These states capture local exploration progress beyond binary visitation. By binding exploration state directly to each visual candidate, MarvisNav enables the VLM to jointly evaluate target relevance and exploration state without a separate post-hoc fusion or reranking stage. Without policy training, MarvisNav achieves state-of-the-art performance on HM3D (81.2% SR and 42.5% SPL), while remaining competitive on MP3D. It also outperforms representative VLM-based methods with far fewer VLM calls (e.g., 7.5% of WMNav). Real-robot experiments across diverse scenes further validate its practical deployability. Beyond MarvisNav, our study shows that memory representation shapes VLM decisions and ZSON performance, highlighting that effective memory use depends not only on its availability, but also on how it is represented. Code and project page will be available at \url{https://wangjincheng1998.github.io/MarvisNav/}.
Comments22 pages