发表机构
School of Computer Science and Information Engineering, Hefei University of Technology; Jianghuai Advance Technology Center; Anhui Provincial Key Laboratory of Humanoid Robots; School of Artificial Intelligence, Beihang University; University of Macau(合肥工业大学计算机与信息工程学院; 江淮先进技术中心; 安徽省人形机器人重点实验室; 北京航空航天大学人工智能学院; 澳门大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究免训练环境下的空中视觉与对话导航任务,提出PSC - AVDN框架,结合解析 - 搜索 - 确认推理管道与结构化空间记忆,整合多种空间线索,在免训练设置中取得新的最优性能。
AI 中文摘要
本文针对资源高效的高空无人机免训练环境下的空中视觉与对话导航(AVDN)任务展开研究。应用大语言模型会因方向定位薄弱和缺乏明确空间线索导致导航不可靠。为此提出PSC - AVDN框架,将三阶段解析 - 搜索 - 确认推理管道与结构化空间记忆紧密结合。解析阶段用大语言模型转换指令,搜索思维链进行目标探索,确认思维链进行精细验证,结构化空间记忆整合多种空间线索。实验表明该框架在免训练环境中建立了新的最优性能。
英文摘要
In this paper, we tackle the Aerial Vision-and-Dialog Navigation (AVDN) task in the training-free setting for resource-efficient high-altitude UAV navigation.Naively applying MLLMs leads to unreliable navigation due to weak directional grounding and the lack of explicit spatial memory.To address these issues, we propose PSC-AVDN, a training-free framework that tightly couples a three-stage Parsing-Search-Confirmation reasoning pipeline with a Structured Spatial Memory (SSM).The parsing stage uses an LLM to convert ambiguous dialogue instructions into stable geometric directional and destination cues.A Search Chain-of-Thought (S-CoT) then performs stepwise target exploration under high-altitude observations, and a Confirmation Chain-of-Thought (C-CoT) conducts fine-grained verification around candidate regions to resolve visual ambiguity.Meanwhile, SSM integrates three complementary sources of spatial cues, including multi-scale visual observation, spatial visual memory, and structured geometric memory to provide global spatial context and long-horizon consistency.Extensive experiments on ANDH and ANDH-Full show that PSC-AVDN establishes new state-of-the-art performance in the training-free setting, matching or surpassing several finetuned methods.Code will be publicly available at: https://github.com/QY6616/PSC-AVDN
Comments10 pages, 4 figures. Accepted to CVPR 2026
Journal refProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2026, pp. 23859-23868