发表机构
D-Robotics; HAI Center, University of Technology Sydney; School of Computer Science, Wuhan University(D-机器人公司; 悉尼科技大学HAI中心; 武汉大学计算机科学学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对VLM智能体零样本对象目标导航存在的问题,提出SkillNav框架,通过分层技能设计及双表示设计,无需训练,在多个数据集上提升成功率,建立新最优成绩。
AI 中文摘要
视觉语言模型(VLM)智能体推动了零样本对象目标导航的发展,但单帧推理使它们缺乏具身导航器所需的跨步骤行为意识,导致出现诸如死胡同停滞、室内循环和迂回接近检测目标等反复出现的失败。基于提示的补救措施会增加跨多子模块情节的令牌预算,并且仍难以编码诸如角度、地图单元格和视点坐标等固有的空间信号。在本文中,我们提出了SkillNav,这是一个基于VLM导航的可扩展行为技能框架,它将现代VLM导航器已经维护的好奇心价值图视为一个可写的基础,在其上可组合技能以零令牌成本铭刻行为记忆。技能根据其行为权限级别分为三层,即用于比例重新加权的软缩放、用于区域级保证的下限提升和用于阈值触发强制动作的硬覆盖,并在固定的组合顺序下跨层协作,该顺序在技能之间建立了可预测的、声明的优先级。这种设计将能力提升转变为技能注册:新行为无需重新训练VLM或干扰现有技能即可插入,为持续改进开辟了道路。一个最小的提示通道用类别级语义提示补充分数级技能,产生一种双表示设计,其中空间记忆存在于地图上,语义记忆存在于短提示中。无需训练即可在MP3D(25.5)、HM3D v0.1(39.3)和HM3D v0.2(43.2)上建立新的最优成功率,相对于最强的先前方法,成功率绝对提高了6.0,并且在HM3D v0.1(69.7)和v0.2(75.9)上实现了最高成功率。
英文摘要
Vision-Language Model (VLM) agents have advanced zero-shot object-goal navigation, yet single-frame reasoning leaves them without the cross-step behavioral awareness an embodied navigator requires, producing recurring failures such as dead-end stalls, in-room loops, and circuitous approaches to detected targets. Prompt-based remedies inflate token budgets across multi-submodule episodes and still struggle to encode inherently spatial signals such as angles, map cells, and viewpoint coordinates. In this paper, we propose SkillNav, an extensible behavioral skill framework for VLM-based navigation that treats the curiosity value map already maintained by modern VLM navigators as a writable substrate on which composable skills inscribe behavioral memory at zero token cost. Skills are stratified into three tiers by their level of behavioral authority, namely soft scaling for proportional reweighting, lower-bound boost for region-level guarantees, and hard override for threshold-triggered forced actions, and cooperate across tiers under a fixed composition order that establishes a predictable, declared priority among skills. This design turns capability improvement into skill registration: new behaviors plug in without retraining the VLM or disturbing existing skills, opening a path for continual refinement. A minimal prompt channel complements the score-level skills with category-level semantic hints, yielding a dual-representation design in which spatial memory lives on the map and semantic memory in short prompts. Training-free, SkillNav establishes new state-of-the-art SPL across MP3D (25.5), HM3D v0.1 (39.3), and HM3D v0.2 (43.2), improving SPL by up to 6.0 absolute over the strongest prior method, and achieves the highest Success Rate on HM3D v0.1 (69.7) and v0.2 (75.9).