VoLN:仅视觉的长距离导航——范式、基准和方法
VoLN: Vision-Only Long-Horizon Navigation---Paradigm, Benchmark, and Method
浏览论文内容
中文总结 AI 辅助
研究提出仅视觉的长距离导航(VoLN)范式,通过VoLN-UAV基准及VoLN-MLLM基线进行实例化。在五个环境测试未见分割中评估,揭示了长距离导航在证据整合、目标匹配和闭环稳定性方面的挑战。
中文摘要 AI 辅助
视觉与语言导航(VLN)使具身智能体能够遵循自然语言指令。然而,在开放的、无GPS的环境中部署时,路线级指令通常编码的空间先验信息,如方向、距离和布局等,无法从机载传感中明确获取。因此,此类接口下的基准性能共同反映了视觉导航能力以及对任务描述中明确提供的路线结构的使用。作为一种补充形式,我们提出了仅视觉的长距离导航(VoLN),它将与路线相关的信息从外部提供的指令和全局引导转移到局部可观察的场景线索中。在VoLN中,目标视图指定目的地,而与路线相关的信息仅通过智能体必须在线检测、解释和选择的局部可观察场景线索来获取。我们通过VoLN-UAV为空中导航实例化VoLN,这是一个包含7210个情节的基准,它结合了长距离目标导向飞行、连续3D运动、大视角变化和上下文相关的信标选择。我们还提供了VoLN-MLLM作为初始参考基线。它将自监督视觉特征与结构化语义空间对齐,并根据观察历史、目标视图、检索到的视觉语义令牌和本体感觉预测短距离航点段。在五个环境的测试未见分割上,它在简单、正常和困难情节上的成功率分别为7.4%、4.5%和1.8%。这些结果对VoLN进行了初步评估,并揭示了在长距离证据整合、跨视图目标匹配和闭环稳定性方面仍存在的重大挑战。项目页面:此https URL
英文摘要
Vision-and-Language Navigation (VLN) enables embodied agents to follow natural-language instructions. However, route-level instructions commonly encode spatial priors, such as orientation, distance, and layout, that are not explicitly available from onboard sensing at deployment in open, GPS-denied environments. Benchmark performance under such interfaces therefore jointly reflects visual navigation ability and the use of route structure explicitly supplied by the task description. As a complementary formulation, we propose Vision-Only Long-Horizon Navigation (VoLN), which shifts route-relevant information from externally supplied instructions and global guidance to locally observable in-scene cues. In VoLN, goal views specify the destination, while route-relevant information is available only through locally observable in-scene cues that the agent must detect, interpret, and select online. We instantiate VoLN for aerial navigation through VoLN-UAV, a 7,210-episode benchmark that combines long-horizon goal-directed flight, continuous 3D motion, large viewpoint changes, and context-dependent beacon selection. We further provide VoLN-MLLM as an initial reference baseline. It aligns self-supervised visual features with a structured semantic space and predicts short-horizon waypoint segments from observation history, goal views, retrieved visual--semantic tokens, and proprioception. On the five-environment Test-Unseen split, it obtains success rates of 7.4%, 4.5%, and 1.8% on Easy, Normal, and Hard episodes, respectively. These results provide an initial evaluation of VoLN and reveal substantial remaining challenges in long-horizon evidence integration, cross-view goal matching, and closed-loop stability. Project page: https://admire-ljb.github.io/VoLN-UAV/
发表机构
- Beihang University(北京航空航天大学)
- Hangzhou International Innovation Institute, Beihang University(北京航空航天大学杭州国际创新研究院)
机构由 AI 辅助整理,请以论文原文为准。