AirGroundVLN:面向目标导向的空地协同视觉与语言导航的大规模基准
AirGroundVLN: A Large-Scale Benchmark for Goal-Oriented Air-Ground Collaborative Vision-and-Language Navigation
浏览论文内容
中文总结 AI 辅助
针对空地协同目标导向VLN缺乏基准和跨平台上下文不一致、空间可观测性不对称的问题,提出含10,281个导航片段的AirGroundVLN基准及AG-CoNAV框架,通过SACM和AGRLP实现有效导航。
中文摘要 AI 辅助
目标导向的视觉与语言导航(VLN)要求智能体在没有预设路线的情况下,根据自然语言描述定位并到达目标。空地协同对于需要广域搜索和细粒度定位的任务具有重要价值。然而,由于缺乏大规模、多样化的基准以及两个核心挑战,目标导向的空地协同VLN的系统性研究仍然受限:1)空中视角与地面视角之间存在显著差异,且随着导航的进行,有用的观测会变得不可用,这使得在跨平台和跨时间维持空间一致的上下文变得困难;2)非对称的空间可观测性使得地面感知在局部细节上丰富但空间范围有限,而空中感知范围广但局部粗糙,这限制了单平台规划的可靠性。为解决这些局限,我们引入了AirGroundVLN,一个包含19个虚幻引擎环境中10,281个导航片段和955个目标实例的基准,并提供了可见/不可见划分以及空中可见性协议以进行系统性评估。伴随该基准,我们提出了AG-CoNAV,一个可训练参考框架,包含两个关键组件:时空锚定协同记忆(SACM)和空中引导的区域到局部规划(AGRLP)。SACM在跨空中和地面观测中维护并检索空间一致的历史上下文,而AGRLP将区域空中引导与细粒度地面导航相结合。大量实验证明了AG-CoNAV的有效性,并将AirGroundVLN确立为未来探索的综合性基准。
英文摘要
Goal-oriented Vision-and-Language Navigation (VLN) requires agents to locate and reach targets described in natural language without prescribed routes. Air--ground collaboration is valuable for tasks requiring both wide-area search and fine-grained localization. However, systematic study of goal-oriented air--ground collaborative VLN remains limited by the lack of large-scale, diverse benchmarks and two core challenges: 1) substantial differences between aerial and ground views, together with useful observations becoming unavailable as navigation proceeds, make it difficult to maintain spatially consistent context across platforms and over time; and 2) asymmetric spatial observability makes ground perception locally detailed but spatially limited and aerial perception broad but locally coarse, limiting the reliability of single-platform planning. To address these limitations, we introduce AirGroundVLN, a benchmark containing 10,281 navigation episodes and 955 target instances across 19 Unreal Engine environments, with seen/unseen splits and an aerial-visibility protocol for systematic evaluation. Alongside the benchmark, we propose AG-CoNAV, a trainable reference framework comprising two key components: Spatiotemporally Anchored Collaborative Memory (SACM) and Aerial-Guided Regional-to-Local Planning (AGRLP). SACM maintains and retrieves spatially consistent historical context across aerial and ground observations. Meanwhile, AGRLP combines regional aerial guidance with fine-grained ground navigation. Extensive experiments demonstrate the effectiveness of AG-CoNAV and establish AirGroundVLN as a comprehensive benchmark for future exploration.
发表机构
- Northwestern Polytechnical University(西北工业大学)
机构由 AI 辅助整理,请以论文原文为准。