发表机构
Shanghai Jiao Tong University; Alibaba Inc; Zhongguancun Academy; Adelaide University; Peking University; Tsinghua University(上海交通大学; 阿里巴巴集团; 中关村学院; 阿德莱德大学; 北京大学; 清华大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出NavMCP框架,结合VLM推理智能体与NFM执行器实现长时程导航,在多个具身问答基准及Unitree Go2机器人上取得优于基线的性能,验证了该方法的有效性。
AI 中文摘要
长时程物理世界智能体必须对遥远目标进行推理,同时将决策建立在可靠的闭环行为基础上。当前的基础模型将这些能力分开:视觉语言模型(VLMs)可推断缺失信息并适配高层计划,但在重复导航的落地方面仍脆弱且低效;而导航基础模型(NFMs)能稳健地执行语义目标,但作为有限的回合运行,缺乏持久的任务级推理。我们提出NavMCP,这是一种智能体式搭建框架,它将VLM推理智能体与NFM执行器结合,用于长时程探索。VLM决定要寻找什么证据、在何处搜索以及何时停止,而NFM将每个语义子目标落地为闭环导航。三个通道构成它们的协作:意图将证据需求转化为导航调用,观测将执行轨迹转化为基于源的轨迹证据,内存则在各次调用间积累发现、负面证据和未解决的目标。该设计将孤立的导航执行轨迹转化为持久的具身交互,无需对任一模型进行重新训练。在具身问答任务中,NavMCP在HM-EQA、MT-HM3D和EXPRESS-Bench上取得了最先进的结果。在匹配智能体和执行器骨干的情况下,它在HM-EQA上比回合式接口高出14.9个百分点。在Unitree Go2机器人上,NavMCP的成功率达到78.3%,且随着任务时长增加,其与最强基线的差距从10个百分点扩大到45个百分点。这些结果表明,将互补的基础模型搭建为长时程物理世界智能体具有潜力。
英文摘要
Long-horizon physical-world agents must reason over distant goals while grounding decisions in reliable closed-loop behavior. Today's foundation models split these capabilities: vision-language models (VLMs) infer missing information and adapt high-level plans but remain brittle and inefficient at repeated navigation grounding, while navigation foundation models (NFMs) robustly execute semantic goals but operate as bounded episodes without persistent task-level reasoning. We introduce NavMCP, an agentic scaffolding framework that couples a VLM reasoning agent with an NFM executor for long-horizon exploration. The VLM decides what evidence to seek, where to search, and when to stop, while the NFM grounds each semantic sub-goal into closed-loop navigation. Three channels structure their collaboration: intent translates evidence needs into navigation calls, observation converts rollouts into source-grounded trajectory evidence, and memory accumulates findings, negative evidence, and unresolved goals across calls. This design turns isolated navigation rollouts into persistent embodied interaction without retraining either model. On Embodied Question Answering, NavMCP achieves state-of-the-art results on HM-EQA, MT-HM3D, and EXPRESS-Bench. Under matched agent and executor backbones, it outperforms an episodic interface by 14.9 percentage points on HM-EQA. On a Unitree Go2, NavMCP reaches 78.3% success, with its margin over the strongest baseline growing from 10 to 45 points as the task horizon increases. These results demonstrate the potential of scaffolding complementary foundation models into long-horizon physical-world agents.
Comments22 pages, 6 figures