PanoVLN:迈向高效的全景视觉-语言导航
PanoVLN: Towards Effective Panoramic Vision-and-Language Navigation
- Zhejiang University(浙江大学)
- The University of Hong Kong(香港大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
PanoVLN通过全景观察、长动作序列预测、分支路线训练及语义几何特征融合,显著提升视觉-语言导航性能,在R2R-CE和RxR-CE上超越SOTA。
AI中文摘要:
近期视觉-语言模型(VLMs)推动了视觉-语言导航(VLN)的发展,使模型能够根据视觉观察和语言指令预测导航动作。在本工作中,我们探索了基于全景观察的VLN,并提出了PanoVLN。动机很直接:更完整的视觉上下文应能带来更明智的导航决策。例如,全景图可以揭示透视相机视野之外的通道,使模型无需额外探索即可识别预期路线。然而,我们发现简单地将透视图像替换为全景图仅带来有限的性能提升。我们的诊断表明,充分利用更广的视野需要对动作预测、训练监督和视觉表示进行修改。首先,更广的视野支持更长视界的动作规划。我们使模型预测更长的动作序列,从而能够从单个全景图实现更大的转向和后续移动。具体而言,我们引入了一种置信度引导执行(CGE)策略,该策略动态决定在重新规划前执行多少个预测动作。其次,更广的视野也带来了更复杂的路线选择。因此,我们构建了具有频繁分支点和清晰指令的训练路线,以提供针对路线选择的有针对性监督。第三,全景导航需要理解跨视角方向的空间关系,而不仅仅是识别单个地标。我们将来自RGB全景图的语义和几何特征相结合,以捕捉场景内容和空间布局,且不增加视觉令牌。凭借4B骨干网络和仅RGB输入,PanoVLN在R2R-CE和RxR-CE Val-Unseen上的成功率分别比之前的SOTA高出11.9%和8.7%。在四足机器人上的真实世界实验进一步表明,与之前的VLN方法相比,导航速度更快且停顿更少。
英文摘要:
Recent vision-language models (VLMs) have advanced vision-and-language navigation (VLN), enabling models to predict navigation actions from visual observations and language instructions. In this work, we explore VLN with panoramic observations and introduce PanoVLN. The motivation is straightforward: more complete visual context should enable better-informed navigation decisions. For example, a panorama can reveal a passage outside a perspective camera's field of view, allowing the model to identify the intended route without additional exploration. However, we find that simply replacing perspective images with panoramas yields only limited gains. Our diagnosis suggests that fully exploiting wider visibility requires modifications to action prediction, training supervision, and visual representation. First, wider visibility supports longer-horizon action planning. We make the model predict longer action sequences, enabling larger turns and subsequent movement from a single panorama. Specifically, we introduce a confidence-guided execution (CGE) strategy that dynamically determines how many predicted actions to execute before replanning. Second, wider visibility also brings more complex route choices. We therefore construct training routes with frequent branching points and clear instructions to provide targeted supervision for route selection. Third, panoramic navigation requires understanding spatial relationships across viewing directions, beyond recognizing individual landmarks. We combine semantic and geometric features from RGB panoramas to capture both scene content and spatial layout without adding visual tokens. With a 4B backbone and RGB-only input, PanoVLN surpasses the previous SOTA by 11.9% and 8.7% in success rate on R2R-CE and RxR-CE Val-Unseen. Real-world experiments on a quadruped further demonstrate faster navigation with fewer pauses than prior VLN methods.