发表机构
Nankai University; Zhongguancun Academy; Shandong University(南开大学; 中关村学院; 山东大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出RAO-Nav,利用全语言模型的隐式视听知识实现零样本语义视听导航,通过潜在导航推理模块增强决策,在公共基准上超越现有方法且无需训练数据。
AI 中文摘要
我们探索全语言模型(OLMs)是否可以直接应用于零样本语义视听导航(SAVN)。近期研究表明,即使是最先进的专业化模型,尽管经过大量任务特定训练,仍难以实现通用多模态导航。在本文中,我们提出了RAO-Nav,即全合一推理OLM(Reasoning All-in-One OLM)的简称,一个用于零样本SAVN的部署流水线。通过利用OLMs中编码的丰富隐式视听知识,具身智能体能够在环境中“听”、“看”、“推理”和“行动”。为了进一步激发OLMs内置的思考能力,我们提出了一种测试时潜在导航推理(LNR)模块,该模块可以无缝集成到解码空间中。LNR鼓励模型检索更多与目标相关的观测并做出有效的导航决策。通过全面的实验,我们展示了我们的框架在公共SAVN基准上超越了现有的最先进基线,且不使用任何训练数据。此外,我们引入了一种新的全局导航指令设置,以进一步评估OLMs作为具身导航代理的能力。代码:此https URL _Nav。
英文摘要
We explore whether Omni-Language Models (OLMs) can be directly applied to zero-shot Semantic Audio-Visual Navigation (SAVN). Recent work demonstrates that even state-of-the-art specialized models still struggle to achieve generalist multimodal navigation, despite extensive task-specific training. In this paper, we introduce RAO-Nav, short for Reasoning All-in-One OLM, a deployment pipeline for zero-shot SAVN. By leveraging the rich implicit audio-visual knowledge encoded in OLMs, the embodied agent is enabled to ``hear'', ``see'', ``reason'', and ``act'' in the environment. To further elicit the built-in thinking ability of OLMs, we propose a test-time Latent Navigation Reasoning (LNR) module that can be seamlessly integrated into the decoding space. LNR encourages the model to retrieve more target-relevant observations and make effective navigation decisions. Through comprehensive experiments, we show that our framework surpasses existing state-of-the-art baselines on public SAVN benchmarks without using any training data. Moreover, we introduce a new \emph{Global Navigation Instruction} setting to further evaluate the ability of OLMs to serve as embodied navigation agents. Code: https://github.com/rikeilong/OmniAV\_Nav.
CommentsAccepted by NeuraIPS 2026