arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AirAnchor:桥接局部与全局空间信息用于零样本航空视觉语言导航

AirAnchor: Bridging Local and Global Spatial Information for Zero-Shot Aerial Vision-and-Language Navigation

Shanwei Fan, Bin Zhang, Zhiwei Xu, Yingxuan Teng, Siqi Dai, Lin Cheng, Guoliang Fan

arXiv 2609.08442首次发表:更新:

发表机构

Institute of Automation, Chinese Academy of Sciences; School of Artificial Intelligence, University of Chinese Academy of Sciences; School of Artificial Intelligence, Shandong University(中国科学院自动化研究所; 中国科学院大学人工智能学院; 山东大学人工智能学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

AirAnchor通过空间锚点桥接局部与全局空间信息,构建统一导航框架,在零样本航空视觉语言导航中显著超越现有基线。

AI 中文摘要

航空视觉语言导航要求无人机遵循自然语言指令并在复杂的城市环境中导航。准确的导航依赖于局部和全局空间信息,它们分别支持即时动作定位和长时程路径规划。然而,现有的零样本方法通常在单一空间尺度上运行,要么依赖于从当前观测在线构建的局部表示,要么依赖于从历史经验离线构建的全局记忆。为解决这一局限,我们提出了AirAnchor,一种通过空间锚点桥接局部与全局空间信息并将其整合到共享导航框架中的新范式,从而为决策提供全面的空间定位。AirAnchor包含三个核心组件:(1)查询驱动的空间锚点定位,从视觉观测中识别与决策相关的锚点并将其组织为局部空间表示;(2)持久对象空间记忆,以增量方式维护一个对象知识库作为持久的全局空间记忆,并检索与地标相关的空间先验;(3)空间信息感知的导航智能体,将局部和全局空间信息显式整合到一个智能体框架中以进行决策。在AerialVLN上的大量实验表明,AirAnchor显著优于现有的零样本基线,验证了所提出范式的有效性和效率。

英文摘要

Aerial Vision-and-Language Navigation requires drones to follow natural-language instructions and navigate through complex urban environments. Accurate navigation relies on both local and global spatial information, which support immediate action grounding and long-horizon path planning, respectively. However, existing zero-shot methods typically operate at a single spatial scale, relying either on local representations constructed online from current observations or on global memories built offline from historical experience. To address this limitation, we propose AirAnchor, a new paradigm that bridges local and global spatial information through spatial anchors and integrates both into a shared navigation framework, enabling comprehensive spatial grounding for decision-making. AirAnchor consists of three core components: (1) Query-Driven Spatial Anchor Grounding, which identifies decision-relevant anchors from visual observations and organizes them into local spatial representations; (2) Persistent Object Spatial Memory, which incrementally maintains an object knowledge base as persistent global spatial memory and retrieves landmark-related spatial priors; and (3) a Spatially-Informed Navigation Agent, which explicitly integrates both local and global spatial information into an agentic framework for decision-making. Extensive experiments on AerialVLN demonstrate that AirAnchor substantially outperforms existing zero-shot baselines, validating the effectiveness and efficiency of the proposed paradigm.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑