发表机构
VinMotion, Inc., Vietnam; University of Southern California, USA(VinMotion公司(越南); 南加州大学(美国))
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
StageVLN通过训练时的空间与轨迹辅助引导,在不增加推理开销的情况下提升视觉语言导航性能,在R2R-CE和RxR-CE上取得显著成果。
AI 中文摘要
视觉语言导航(VLN)策略日益受益于大型视觉语言模型(VLM)提供的强语义先验。然而,标准的动作监督并未明确鼓励中间表示保留场景几何、相对朝向或全局情节进度。在推理时引入深度估计器、显式地图、点云或几何基础模型可以提供此类结构,但会带来额外的计算量、内存开销和部署时的架构依赖。我们提出StageVLN,一种通过特权空间与轨迹引导来塑造导航表示、同时保留原始推理路径的训练框架。一个冻结的几何基础模型为分层导航器状态提供多级空间引导,而相对航向和专家路线进度目标提供互补的轨迹状态监督。所有辅助组件仅在训练期间使用,部署时移除。在R2R-CE验证未见环境中,StageVLN在4B参数骨干下达到56.3%的SR和51.4%的SPL,推理时无需额外几何编码器。在RxR-CE上,它达到54.3%的SR,无需额外导航训练数据或推理时的几何编码器。
英文摘要
Vision-and-Language Navigation (VLN) policies increasingly benefit from strong semantic priors provided by large vision-language models (VLMs). However, standard action supervision does not explicitly encourage intermediate representations to preserve scene geometry, relative orientation, or global episode progress. Incorporating depth estimators, explicit maps, point clouds, or geometry foundation models at inference can provide such structure but introduces additional computation, memory overhead, and architectural dependence during deployment. We introduce StageVLN, a training framework that shapes navigation representations through privileged spatial and trajectory guidance while preserving the original inference pathway. A frozen geometry foundation model provides multi-level spatial guidance to hierarchical navigator states, while relative-heading and expert-route progress objectives provide complementary trajectory-state supervision. All auxiliary components are used only during training and removed at deployment. On R2R-CE validation-unseen, StageVLN achieves 56.3\% SR and 51.4\% SPL with a 4B-parameter backbone, without an additional geometry encoder at inference. On RxR-CE, it achieves 54.3\% SR without additional navigation training data or a geometry encoder at inference.