arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.10383cs.CVcs.AIcs.RO

ABot-N1:迈向通用视觉语言导航基础模型

ABot-N1: Toward a General Visual Language Navigation Foundation Model

Ruiyan Gong, Yingnan Guo, Junjun Hu, Jintao Kong, Xiaoxu Leng, Tianlun Li, Weize Li, Fei Liu, Zhicheng Liu, Jia Lu, Minghua Luo, Chenlin Ming, Yanfen Shen, Jiyu… 展开作者

Ruiyan Gong, Yingnan Guo, Junjun Hu, Jintao Kong, Xiaoxu Leng, Tianlun Li, Weize Li, Fei Liu, Zhicheng Liu, Jia Lu, Minghua Luo, Chenlin Ming, Yanfen Shen, Jiyue Tao, Zhengbo Wang, Mingyang Yin, Minqi Gu, Zihao Guan, Wei Guo, Guoqing Liu, Huachong Pang, Menglin Yang, Zeqian Ye, Xiaoxiao Geng, Zhining Gu, Honglin Han, Di Jing, Hongyu Pan, Mingchao Sun, Kuan Yang, Jianfang Zhang, Yanghong Chen, Ye He, Wei Mei, Jiahao Shi, Xiangpo Yang, Yanqing Zhu, Yang Cai, Jingjing Ma, Shihui Su, Zixiao Tang, Linbo Zheng, Zedong Chu, Xiaolong Wu, Wenbin Tang, Mu Xu

首次发表
浏览论文内容

中文总结 AI 辅助

研究旨在构建通用视觉语言导航基础模型,针对当前方法问题,提出ABot-N1,通过慢-快架构解耦认知与控制,利用双视觉语言信号,实现跨场景稳健、可推广且可解释的导航,在城市规模导航等任务中表现出色并开源新基准。

中文摘要 AI 辅助

视觉语言导航基础模型旨在统一深度推理,实现有根据的空间决策,并具备广泛的通用性以适用于各种具身任务。当前方法通常通过将观察直接映射到动作的整体策略来实现这种整合,但存在坐标漂移和长尾语义处理不佳的问题,且缺乏可解释性。本文提出ABot-N1,通过慢-快架构将认知与控制解耦,由双视觉语言信号引导,解决了这些挑战。慢视觉语言推理器在生成像素目标时进行显式的思维链推理,为多种任务提供通用接口。快速动作专家利用文本线索和像素指导生成连续航点。该方法在模拟和真实世界基准测试中确保了强大、可推广且可解释的导航。ABot-N1建立了新的技术记录,在城市规模导航中尤其有显著提升,还在其他任务中保持了卓越的鲁棒性。新的点目标/兴趣点目标基准作为开源发布,以推动城市规模导航领域的发展。

英文摘要

Visual Language Navigation foundation models aim to unify deep reasoning for grounded spatial decisions with broad versatility for diverse embodied tasks. Current approaches typically achieve this integration via monolithic policies that map observations directly to actions, yet they often suffer from coordinate drift and poor handling of long-tail semantics. Furthermore, these black-box mappings lack interpretability, hindering the simultaneous achievement of generality, robustness, and transparency. We present ABot-N1, a step toward a general Visual Language Navigation foundation model, that addresses these challenges by decoupling cognition from control via a slow-fast architecture guided by dual visual-language signals. More specifically, a slow vision-language reasoner performs explicit Chain-of-Thought reasoning while producing a pixel goal. This compact set of image-space anchor points serves as a universal interface for diverse tasks, including point-goal, object-goal, poi-goal, instruction-following, and person-following. Subsequently, a fast action expert leverages both the textual cues and the pixel guidance to generate continuous waypoints at the native control frequency. By bridging high-level intents and low-level control through pixel-grounded anchors paired with explicit linguistic traces, our approach ensures robust, generalizable, and interpretable navigation across simulation and real-world benchmarks. ABot-N1 establishes new state-of-the-art records, delivering massive gains specifically in urban-scale navigation: boosting POI arrival by 35.0% (to 77.3%) and achieving 95.4%/92.9% SR in complex indoor and outdoor scenes. It also maintains superior robustness across object-reaching, person-following, and instruction-following tasks. New Point-Goal/POI-Goal benchmarks are released as open source to advance the field of urban-scale navigation.

↑