arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2504.21432cs.ROcs.CV

UAV-VLN:面向无人机的端到端视觉语言引导导航

UAV-VLN: End-to-End Vision Language guided Navigation for UAVs

Pranav Saxena, Nishant Raghuvanshi, Neena Goveas

首次发表 更新
浏览论文内容

中文总结 AI 辅助

本文提出端到端UAV-VLN框架,融合LLM常识推理与视觉目标定位,通过跨模态落实实现无人机依据自然语言指令在陌生室内外环境中高效导航。

中文摘要 AI 辅助

AI引导自主系统中的一个核心挑战,是使智能体能够根据自然语言命令,在先前未见过的环境中真实且有效地导航。我们提出 UAV-VLN,这是一种面向无人机(Unmanned Aerial Vehicles,UAVs)的新型端到端视觉语言导航(Vision-Language Navigation,VLN)框架,可将大语言模型(Large Language Models,LLMs)与视觉感知无缝集成,以促进人机交互式导航。我们的系统解释自由形式的自然语言指令,将其落实到视觉观测中,并在多样化环境中规划可行的空中轨迹。UAV-VLN利用LLMs的常识推理能力解析高层语义目标,同时由视觉模型检测并定位环境中与语义相关的对象。通过融合这些模态,无人机能够推理空间关系、消除人类指令中指称的歧义,并在最少的任务特定监督下规划具备情境感知能力的行为。为确保稳健且可解释的决策,该框架包含一个跨模态落实机制,用于将语言意图与视觉情境对齐。我们在多样化的室内和室外导航场景中评估了UAV-VLN,证明其能够以最少的任务特定训练泛化到新的指令和环境。我们的结果显示,指令遵循准确率和轨迹效率均取得显著提升,突显了由LLM驱动的视觉语言接口在安全、直观且可泛化的无人机自主方面的潜力。

英文摘要

A core challenge in AI-guided autonomy is enabling agents to navigate realistically and effectively in previously unseen environments based on natural language commands. We propose UAV-VLN, a novel end-to-end Vision-Language Navigation (VLN) framework for Unmanned Aerial Vehicles (UAVs) that seamlessly integrates Large Language Models (LLMs) with visual perception to facilitate human-interactive navigation. Our system interprets free-form natural language instructions, grounds them into visual observations, and plans feasible aerial trajectories in diverse environments. UAV-VLN leverages the common-sense reasoning capabilities of LLMs to parse high-level semantic goals, while a vision model detects and localizes semantically relevant objects in the environment. By fusing these modalities, the UAV can reason about spatial relationships, disambiguate references in human instructions, and plan context-aware behaviors with minimal task-specific supervision. To ensure robust and interpretable decision-making, the framework includes a cross-modal grounding mechanism that aligns linguistic intent with visual context. We evaluate UAV-VLN across diverse indoor and outdoor navigation scenarios, demonstrating its ability to generalize to novel instructions and environments with minimal task-specific training. Our results show significant improvements in instruction-following accuracy and trajectory efficiency, highlighting the potential of LLM-driven vision-language interfaces for safe, intuitive, and generalizable UAV autonomy.

发表机构

  • Birla Institute of Technology and Science Pilani, K.K Birla Goa Campus(比拉理工学院和科学学院,K.K比拉果阿校区)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑