C$^2$Nav:先比较再承诺——用于零样本视觉与语言导航
C$^2$Nav: Compare Before You Commit for Zero-Shot Vision-and-Language Navigation
- North China University of Technology(北方工业大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
C2Nav提出先比较再承诺的VLM-机器人接口,通过序数比较替代基数输出,在零样本VLN-CE中无需训练即达44%SR,证明约束接口与强推理互补。
AI中文摘要:
连续环境中的零样本视觉与语言导航(VLN-CE)日益将基础视觉语言模型(VLM)置于导航回路中。现有系统通常要求输出基数型结果,如航路点、像素、朝向、进度值或绝对到达决策,从而将生成式响应与几何量级或不可逆的承诺耦合在一起。我们研究了一种互补的模型-机器人接口:VLM对控制器构建的备选方案进行比较,而几何、阈值、动作量级和执行则保留在物理侧。我们将这一思想实例化为C2Nav,一个无需训练、具有三个协调能力的框架。“观察”对经过物理验证的候选视图执行序数型视线选举;“记忆”维护一条紧凑的路线草图并比较相邻指令片段假设;“到达”结合犹豫阶梯、回看比较和可撤销的回退行走以实现可靠的停止。在公开的OpenNav R2R-CE 100协议上,使用Qwen3-VL-8B-Instruct的C2Nav获得41.0%的OSR、31.0%的SR和16.7%的SPL,而使用标准GPT-5.5模型的同一接口达到54.0%的OSR、44.0%的SR和29.0%的SPL。全能力消融实验显示,缺少“观察”时SR降至14.0%,缺少“记忆”时降至25.0%,缺少“到达”时降至29.0%。仅将比较式回答形式替换为基数/绝对问题的匹配角色反转,在空间、过渡和终止槽位中分别将SR降至12.0%、28.0%和21.0%。结果表明,受约束的决策接口与更强的VLM推理是互补的,而非可互换的。
英文摘要:
Zero-shot vision-and-language navigation in continuous environments (VLN-CE) increasingly places foundation vision-language models (VLMs) inside the navigation loop. Existing systems commonly request cardinal outputs such as waypoints, pixels, headings, progress values, or absolute arrival decisions, coupling a generative response to geometric magnitude or an irreversible commitment. We study a complementary model-robot interface: the VLM compares controller-constructed alternatives, while geometry, thresholds, action magnitude, and execution remain on the physical side. We instantiate this idea in C2Nav, a training-free framework with three coordinated faculties. Seeing performs ordinal Gaze Election over physically vetted candidate views; Remembering maintains a compact route sketch and compares adjacent instruction-leg hypotheses; and Arriving combines a hesitation ladder, look-back comparison, and revocable walk-back for reliable stopping. On the public OpenNav R2R-CE 100 protocol, C2Nav with Qwen3-VL-8B-Instruct obtains 41.0% OSR, 31.0% SR, and 16.7% SPL, while the same interface with the standard GPT-5.5 model reaches 54.0% OSR, 44.0% SR, and 29.0% SPL. Whole-faculty ablations reduce SR to 14.0% without Seeing, 25.0% without Remembering, and 29.0% without Arriving. Matched role inversions that replace only the comparative answer form with cardinal/absolute questions reduce SR to 12.0%, 28.0%, and 21.0% in the spatial, transition, and terminal slots, respectively. The results indicate that a constrained decision interface and stronger VLM reasoning are complementary rather than interchangeable.