arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AirForesight:面向无人机视觉语言导航(UAV-VLN)的跨空间规划一致性当前-未来空间地图想象

AirForesight: Current-to-Future Spatial Map Imagination with Cross-Space Planning Consistency for UAV-VLN

Yutong Liu, Xiaojie Li, Mingzhu Xu, Jianlong Wu

arXiv 2608.12835首次发表:更新:

发表机构

Harbin Institute of Technology, Shenzhen; Shandong University(哈尔滨工业大学(深圳); 山东大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对现有UAV-VLN方法空间推理不足的问题,提出AirForesight框架,通过联合监督学习当前地图表示并引入跨空间规划一致性损失,在公开数据集上取得良好性能。

AI 中文摘要

无人机视觉语言导航(UAV-VLN)要求智能体遵循语言指令,从稀疏多视图观测中推断空间结构,并在复杂室外环境中执行可行的三维运动。尽管大型语言模型已取得近期进展,但多数现有方法仍将视觉-语言输入直接映射为动作,提供的显式场景 grounding 和感知未来的空间推理能力有限。本文提出面向UAV-VLN的当前-未来空间地图想象框架AirForesight,该框架首先从多视图观测中学习结构化当前地图表示,该表示受当前地图重构与未来轨迹预测的联合监督,以编码当前场景结构与未来运动意图。在结构化因果注意力机制下,当前空间知识被传播至未来地图推理,聚合得到的当前与未来表示用于预测下一个三维航路点。为使空间想象更贴合导航需求,本文引入跨空间规划一致性损失,该损失鼓励预测的地图空间轨迹与真实航路点位移导出的专家动作方向保持方向一致性。在OpenUAV和AerialVLN-S数据集上的实验及大量 ablation 研究表明,所提框架性能强劲,且其有效性与稳定性得到验证。

英文摘要

Unmanned Aerial Vehicle Vision-Language Navigation (UAV-VLN) requires agents to follow language instructions, infer spatial structure from sparse multi-view observations, and execute feasible 3D motion in complex outdoor environments. Despite recent progress with large language models, most existing methods still map vision-language inputs directly to actions, providing limited explicit scene grounding and future-aware spatial reasoning. We propose AirForesight, a current-to-future spatial map imagination framework for UAV-VLN. AirForesight first learns a structured current-map representation from multi-view observations. This representation is jointly supervised by current-map reconstruction and future-trajectory prediction, encouraging it to encode both present scene structure and future motion intent. Under structured causal attention, the current spatial knowledge is propagated to future-map reasoning, and the resulting current and future representations are aggregated to predict the next 3D waypoint. To make spatial imagination more relevant to navigation, we introduce a cross-space planning consistency loss that encourages directional agreement between the predicted map-space trajectory and the expert action direction derived from the ground-truth waypoint displacement. Experiments on OpenUAV and AerialVLN-S, together with extensive ablations, demonstrate strong performance and support the effectiveness and stability of the proposed framework.

CommentsAccepted by ACM Multimedia 2026

DOI:10.1145/3767308.3836065

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑