arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

CoNav-UAV:基于Stackelberg学习的协同双高度无人机导航

CoNav-UAV: Cooperative Dual-Altitude Aerial Navigation via Stackelberg Learning

Junru Song, Wenhao Zhang, Yang Yang, Xuekai Qiu, Feifei Wang, Weien Zhou, Tingsong Jiang, Ying Wen, Yang Li, Wen Yao

arXiv 2608.01802首次发表:更新:

AI 中文总结

针对双高度无人机协同导航的合作问题,本文提出CoNav-UAV,通过迭代Stackelberg学习优化高空领导者与低空跟随者,在AerialVLN基准上显著提升导航成功率且适配数据需求更少。

AI 中文摘要

面向空中平台的目标导向视觉与语言导航(VLN)因灾难救援、基础设施巡检、安全巡逻等任务受到日益关注。该任务中,无人机(UAV)仅依靠目标外观与周边环境的简洁描述完成目标定位,需兼顾全局探索、环境感知及无碰撞近距离接近,这两个交织的过程难以在单个智能体中协调。现有多数方法将地面VLN范式迁移至低空无人机,通过外部辅助弥补其探索效率低的问题;近期有研究部署两架互补高度的无人机,但仍依赖特权信息且两个智能体独立训练,缺乏合作必需的相互适应。本文提出CoNav-UAV,将任务明确建模为高空领导者与低空跟随者间的Stackelberg博弈,系统仅基于机载视觉与语言输入运行。为求解该博弈,引入迭代Stackelberg学习:领导者的高层视觉语言推理通过基于记忆的上下文学习优化,跟随者的精确运动控制通过DAgger式专家蒸馏更新,二者交替优化直至达到Stackelberg均衡。在AerialVLN基准的三个高保真城市场景中,CoNav-UAV始终优于单智能体与双智能体基线:学习场景的成功率提升最高达30.8个百分点,跨场景迁移时提升9.0个百分点,且仅需约1/3的适配数据;进一步分析验证了领导者与跟随者更新的互补增益,还揭示其在不同视觉语言模型(VLM)骨干上具备稳健增益但学习动态存在差异。

英文摘要

Target-oriented vision-and-language navigation (VLN) on aerial platforms is attracting growing attention for missions such as disaster rescue, infrastructure inspection, and security patrol. In this task, an unmanned aerial vehicle (UAV) needs to locate targets given only a concise description of their appearance and surroundings. This requires global exploration and grounding as well as collision-free close-range approach, two interleaved processes difficult to reconcile within a single agent. Most existing methods transfer the ground VLN paradigm to a low-altitude UAV and compensate for its inefficient exploration with external assistance. A recent attempt deploys two UAVs at complementary altitudes yet still relies on privileged information and trains its two agents independently, precluding any mutual adaptation essential for cooperation. Here we propose CoNav-UAV, which explicitly models the task as a Stackelberg game between a high-altitude leader and a low-altitude follower, with the system operating on onboard visual and linguistic inputs alone. To solve this game, we introduce Iterative Stackelberg Learning. The leader's high-level vision-language reasoning is refined via memory-based in-context learning, while the follower's precise motion control is updated via DAgger-style expert distillation. The alternation drives both agents toward a Stackelberg equilibrium. CoNav-UAV consistently outperforms single- and dual-agent baselines across three high-fidelity urban scenes from the AerialVLN benchmark. Success rate improves by up to 30.8 points on the learning scene, and 9.0 points under cross-scene transfer while using about 3x less adaptation data. Further analyses validate the complementary gains of the leader and follower updates and reveal robust gains yet distinct learning dynamics across VLM backbones.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑