发表机构
Uni-Ubi AI; Zhejiang University; Tongji University(优必爱人工智能; 浙江大学; 同济大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出面向目标的视频 grounding 导航指令生成任务VideoNIG,设计两阶段课程学习框架解决该任务,实验表明其可提升指令质量,集成VLN智能体后可实现端到端导航。
AI 中文摘要
导航指令生成(Navigation Instruction Generation, NIG)旨在生成用于导航引导的逐步自然语言指令。现有研究主要将NIG视为视觉语言导航(Vision-and-Language Navigation, VLN)的辅助任务,聚焦于数据增强或多任务学习。然而,从紧凑的环境先验生成导航指令需要细致的空间推理,尤其是当目标路线并非简单跟随已演示的游览路线时,这对当前多模态模型仍是挑战。本研究提出VideoNIG,这是一种面向目标的视频 grounding NIG任务,可从以自我为中心的游览视频、初始观测以及文本或视觉目标生成导航指令,无需依赖图、地图等中间表示。我们在受控模拟器基准中实例化VideoNIG,涵盖连续室内环境中的60000个游览视频和具有递进难度级别的37000个多模态提示。我们进一步引入诊断评估协议,结合文本相似度、基于选择的空间一致性测试以及下游导航执行。为解决该任务,我们提出两阶段课程学习框架,将学习分解为基础动作感知和长程导航推理。具体而言,我们首先采用动作热身(Action Warmup)实现空间动作-视角对齐,随后采用复杂度递进(Complexity Progression),使用探索难度递增的轨迹。大量实验表明,现有多模态大语言模型(MLLM)在VideoNIG任务中表现不佳,而我们的方法在互补诊断指标上显著提升了指令质量。最后,将VideoNIG生成的指令与VLN智能体集成,证明了该任务形式化用于端到端导航的可执行性。
英文摘要
Navigation Instruction Generation (NIG) aims to produce step-by-step natural language instructions for navigation guidance. Existing studies primarily treat NIG as an auxiliary task for vision-andlanguage navigation (VLN), focusing on data augmentation or multi-task learning. However, generating navigation instructions from compact environmental priors requires meticulous spatial reasoning, especially when the target route does not simply follow the demonstrated tour, and remains challenging for current multimodal models. In this work, we introduce VideoNIG, a goal-oriented video-grounded NIG task that generates navigation instructions from ego-centric tour videos, an initial observation, and a textual or visual goal, without relying on intermediate representations such as graphs and maps. We instantiate VideoNIG in a controlled simulator benchmark with 60K tour videos across continuous indoor environments and 37K multimodal prompts with progressive difficulty levels. We further introduce a diagnostic evaluation protocol that combines text similarity, choice-based spatial consistency tests, and downstream navigation execution. To address this task, we propose a two-stage Curriculum Learning framework that decomposes the learning into foundational motion perception and long-horizon navigation reasoning. Specifically, we first employ Action Warmup for spatial action-view alignment, followed by Complexity Progression using trajectories with increasing exploratory difficulty. Extensive experiments show that existing MLLMs struggle with VideoNIG, while our approach significantly improves instruction quality across complementary diagnostic metrics. Finally, integrating VideoNIG-generated instructions with a VLN agent demonstrates the executability of this task formulation for end-to-end navigation.