发表机构
State Key Laboratory of Virtual Reality Technology and Systems, Beihang University; School of Computer Science and Engineering, Beihang University(北京航空航天大学虚拟现实技术与系统国家重点实验室; 北京航空航天大学计算机科学与工程学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文针对无人机视觉语言导航的现有问题,提出统一语义到决策框架,经实验在AerialVLN和OpenFly基准上达到SOTA性能。
AI 中文摘要
无人机视觉语言导航(UAV-VLN)旨在使空中智能体能在开放3D环境中,基于自中心视觉观测遵循自然语言指令。现有方法存在三个耦合问题:视觉观测中与指令相关的地标接地性弱、长 horizon 历史信息利用不足、局部陷阱或重复探索下决策不稳定。为解决这些问题,本文提出统一的语义到决策框架:首先,设计指令接地语义增强模块,将物体级语义与相对空间线索注入当前观测状态;其次,开发关联感知动态时间聚合策略,对完整历史缓冲区重新加权,同时将少数高关联帧转换为结构化地标提示输入解码器;最后,提出拓扑感知决策方法,结合局部最优认知与组相对策略优化,基于进度、目标、语义及路径合规性奖励执行。在广泛使用的AerialVLN与OpenFly基准上的实验清晰表明,本文方法实现了SOTA性能。
英文摘要
UAV vision-language navigation (UAV-VLN) focuses on enabling an aerial agent to follow natural-language instructions in open 3D environments from egocentric visual observations. Current approaches suffer from three coupled issues: weak grounding of instruction-relevant landmarks in visual observations, insufficient exploitation of long-horizon history, and unstable decisions under local traps or repeated exploration. To address these issues, we propose a unified semantic-to-decision framework. First, we present an instruction-grounded semantic enhancement module that injects object-level semantics and relative spatial cues into the current observation state. Subsequently, we develop a relevance-aware dynamic temporal aggregation strategy that reweights the full history buffer while converting a few high-relevance frames into structured landmark prompts for the decoder. Finally, we devise a topology-aware decision method that combines local-optimum cognition with group-relative policy optimization under progress, goal, semantic, and path-compliance rewards. Experiments on the widely used AerialVLN and OpenFly benchmarks clearly demonstrate that our method achieves state-of-the-art performance.
Comments10 pages, 5 figures. Accepted at ACM Multimedia 2026 (MM '26)