从路线到步骤:在视觉语言导航中分离语义进展与局部执行
From Routes to Steps: Separating Semantic Progress from Local Execution in Vision-and-Language Navigation
AI总结:
本研究针对视觉语言导航中进展与执行错误难区分的问题,提出Route2Step框架,通过步骤级接口分离语义进展跟踪与动作生成,在R2R-CE数据集上显著提升了导航性能,且在真实环境中具备实用性。
AI中文摘要:
视觉语言导航(VLN)要求智能体遵循路线级指令,从自我中心视觉观测中执行其组成步骤。现有基于视觉语言模型(VLM)的导航器通常仅通过下一步动作预测来监督这两种能力,导致进展跟踪错误与执行错误难以区分。当智能体偏离路线时,纠正动作标签可能恢复下一个动作,但无法表明智能体是选择了错误的子指令,还是未能执行正确的子指令。因此,智能体可能从错误的进展状态继续做出决策。为解决这种歧义,我们提出Route2Step框架,该框架通过显式步骤级接口将语义进展跟踪与动作生成分离开来。指令分析模块($\boldsymbol{\textit{M}}_{\text{IA}}$)根据全局指令和视觉历史预测该状态。动作生成模块($\boldsymbol{\textit{M}}_{\text{AG}}$)以预测的状态和近期观测为条件,生成本地动作块。为在无需手动时间标签的情况下监督进展状态,我们提出E-SPA(一种步骤对齐过程),将子指令与路线级演示的对应部分相关联。这些对齐方式支持对错误进展估计的状态监督,而直接动作监督则保留给在正确激活子指令下反复失败的部署组。在R2R-CE数据集上,Route2Step使用19万个状态级纠正样本,仅需1.15万个直接动作监督状态,就将成功率(SR)从48.1%提升至55.3%,将路径预测成功率(SPL)从43.3%提升至48.2%。在真实室内和室外环境中的实验进一步证明了Route2Step的实际适用性。项目页面为:this https URL。
英文摘要:
Vision-and-Language Navigation (VLN) requires an agent to follow a route-level instruction by executing its constituent steps from egocentric visual observations. Existing VLM-based navigators typically supervise both capabilities through next-action prediction alone, making progress-tracking errors difficult to distinguish from execution errors. When an agent deviates from the route, a corrective action label may recover the next movement but does not indicate whether the agent selected the wrong sub-instruction or failed to execute the correct one. Consequently, the agent may continue making decisions from an erroneous progress state. To resolve this ambiguity, we propose \textbf{Route2Step}, a framework that decouples semantic progress tracking from action generation through an explicit step-level interface. The Instruction Analysis Module ($\mathcal{M}_{\mathrm{IA}}$) predicts this state from the global instruction and visual history. Conditioned on the predicted state and recent observations, the Action Generation Module ($\mathcal{M}_{\mathrm{AG}}$) generates local action chunks. To supervise the progress state without manual temporal labels, E-SPA, a step-alignment procedure, associates sub-instructions with their corresponding portions of route-level demonstrations. These alignments enable state supervision for incorrect progress estimates, while direct action supervision is reserved for rollout groups that repeatedly fail under the correct active sub-instruction. On R2R-CE, Route2Step improves SR from 48.1\% to 55.3\% and SPL from 43.3\% to 48.2\%, using 190K state-level corrective samples while requiring only 11.5K directly action-supervised states. Experiments in real-world indoor and outdoor environments further demonstrate the practical applicability of Route2Step. The project page is: https://sisyphus-hxy.github.io/Route2Step/.