VTInstructor:面向连续环境中导航指令生成的视觉轨迹提示
VTInstructor: Visual Trajectory Prompting for Navigation Instruction Generation in Continuous Environments
- Peking University(北京大学)
- PrimeBot
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本研究提出首个面向连续环境的VLN指令生成框架VTInstructor,通过视觉轨迹提示技术生成导航指令,在基准测试中刷新最优性能,提升跟随器成功率并为下游任务带来数据增强增益。
AI中文摘要:
从连续环境下的第一人称视角RGB视频生成导航指令,是人机交互与可扩展数据集构建领域一项重要但颇具挑战性的任务。现有指令生成器均基于带有全景观测的离散视点图,其中轨迹结构明确;但在连续环境中,智能体仅能接收密集的RGB流,难以恢复轨迹线索。我们提出VTInstructor,首个面向连续环境的视觉语言导航(VLN)指令生成框架。其核心思路是将隐式轨迹几何转换为显式视觉轨迹提示:EDTC将长RGB轨迹浓缩为导航关键关键帧,VTP将路径、转向和目标线索叠加到这些锚点上,VTMod将所得轨迹信号注入视觉编码器,VT-GRPO在训练期间进一步校准该空间注入,全程无需导航图、预建地图或场景重建。在极具挑战性的R2R-CE与RxR-CE验证未见过基准上,VTInstructor在所有标准自然语言生成(NLG)指标上创下新的最优性能,分别超越最强基线0.357 CIDEr与0.109 CIDEr。除自动指标外,VTInstructor生成的指令将冻结跟随器的成功率提升至63.3%,较最优竞争指令源提升14.7个百分点,还为下游导航任务带来+3个SR点的一致数据增强增益。
英文摘要:
Navigation instruction generation from ego-centric RGB video in continuous environments is an important yet challenging task for human-robot interaction and scalable dataset construction. Prior instruction generators assume discrete viewpoint graphs with panoramic observations, where trajectory structure is explicit; in continuous environments, however, the agent receives only a dense RGB stream, making trajectory cues difficult to recover. We propose VTInstructor, the first VLN instruction generation framework for continuous environments. Our key idea is to convert implicit trajectory geometry into explicit visual trajectory prompts: EDTC condenses long RGB trajectories into navigation-critical keyframes, VTP overlays path, turn, and goal cues onto these anchors, VTMod injects the resulting trajectory signals into the visual encoder, and VT-GRPO further calibrates this spatial injection during training, all without requiring a navigation graph, pre-built map, or scene reconstruction. On the challenging R2R-CE and RxR-CE Val Unseen benchmarks, VTInstructor sets a new state of the art across all standard NLG metrics, surpassing the strongest baseline by +0.357 CIDEr and +0.109 CIDEr, respectively. Beyond automatic metrics, VTInstructor-generated instructions raise a frozen follower's success rate to 63.3%, a +14.7 percentage-point gain over the best competing instruction source, and provide consistent data augmentation gains of +3 SR points on downstream navigation tasks.