arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

NavGPT-3:在分层导航运行时中利用上下文

NavGPT-3: Harnessing Context in a Hierarchical Navigation Runtime

Gengze Zhou, Yicong Hong, Jiazhao Zhang, Xunyi Zhao, Jian Zhou, Zixing Lei, Zun Wang, Chongyang Zhao, Xionghui Chen, Stephen Gould, Anton van den Hengel, Qi Wu

arXiv 2610.10787首次发表:更新:

发表机构

Adelaide University; Roblox; PKU; SJTU; UNC Chapel Hill; UNSW; Metacognition; ANU(阿德莱德大学; 罗布乐思公司; 北京大学; 上海交通大学; 北卡罗来纳大学教堂山分校; 新南威尔士大学; 元认知公司; 澳大利亚国立大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出NavGPT-3,通过类操作系统运行时连接语言模型与动作策略模型,在R2R-CE和RxR-CE上达到最优性能,首次使自主智能体在RxR-CE上达到人类水平,将发布相关模型与代码。

AI 中文摘要

通过长程智能体强化学习训练的语言模型能够通过推理泛化知识、表达精确动作并在多步骤中追求目标,提升了具身智能体的理解与决策上限。然而,物理交互仍属于动作策略的范畴,其提供密集、低延迟的控制。本文提出NavGPT-3,这是一个连接上述两类模型的工具,在其之上构建了类操作系统的运行时:推理、动作和监控作为线程运行,拥有各自的上下文、工具和权限,运行时对这些线程进行调度并决定哪个线程控制机器人的运动,使机器人能通过中断和线程切换对突发现实世界事件做出反应。在底层,我们的动作策略模型NavGPT VLA在1928万样本上训练,通过编解码器分配按场景变化比例分配视觉令牌;其8B模型在R2R-CE上达到74.51的成功率(SR),在RxR-CE上以78.19的SR领先。借助完整的工具链,NavGPT-3在R2R-CE上达到81.51的SR,首次使自主智能体达到人类水平:在RxR-CE上,其成功度(90.43 vs. 90.4 SR)和路径保真度(78.47 vs. 77.7 nDTW)与人类跟随者相当,每集耗时1分22秒,而人类约需3分钟。我们对工具链设计及两类模型的交互进行了全面消融实验,揭示工具和动作策略如何塑造从语言模型推理到物理控制的路径:当NavGPT VLA执行路线时,推理循环缩短,系统的最小反应时间从语言模型决策时的3-19秒降至动作策略步骤的0.5-1秒(1-2赫兹)。这些结果表明,设计这类具身接口是连接前沿语言模型智能与低层物理控制的关键。我们将发布所有模型、代码和评估记录。

英文摘要

Language models trained with long-horizon agentic reinforcement learning can generalize knowledge through reasoning, express precise actions, and pursue goals over many steps, raising the ceiling on what an embodied agent can understand and decide. Physical interaction, however, remains the domain of action policies, which provide dense, low-latency control. We present NavGPT-3, a harness that connects the two models, with an OS-like runtime built above it: reasoning, acting, and monitoring run as threads with their own context, tools, and permissions, while the runtime schedules them and decides which thread controls the robot's motion, so that the robot can react to sudden real-world events through interruption and thread switching. Beneath it, our action policy NavGPT VLA, trained on 19.28M examples, allocates visual tokens using codec allocation, in proportion to scene change; its 8B model alone reaches 74.51 SR on R2R-CE and leads RxR-CE with 78.19 SR. With the complete harness, NavGPT-3 sets the state of the art on R2R-CE (81.51 SR) and, for the first time, brings an autonomous agent to human level: on RxR-CE it matches human followers in success (90.43 vs. 90.4 SR) and path fidelity (78.47 vs. 77.7 nDTW) at 1 min 22 s per episode, versus roughly 3 min for a human. We comprehensively ablate the harness design and the interaction between the two models, showing how tools and the action policy shape the path from language-model reasoning to physical control: when NavGPT VLA executes the route, the reasoning loop shortens and the system's minimum reaction time falls from 3-19 s per language-model decision to 0.5-1 s per action-policy step (1-2 Hz). These results show that designing this embodied interface is central to connecting frontier language-model intelligence with low-level physical control. We will release all models, code, and evaluation records.

Comments36 pages, 14 figures. Project page: https://metacognitionai.github.io/NavGPT3/

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑