基于共享鸟瞰图的空地协同视觉语言导航
Air-Ground Collaborative Vision-and-Language Navigation via Shared Bird's-Eye Maps
浏览论文内容
中文总结 AI 辅助
本文提出首个空地协同视觉语言导航的无训练基线AGC-VLN,通过共享鸟瞰图实现无人机与地面无人车的协同,在CARLA-Air场景中联合成功率达77.0%,显著优于单智能体基线。
中文摘要 AI 辅助
空地协同视觉语言导航(VLN)将具有全局鸟瞰视角的无人机(UAV)与具有局部第一人称视角的地面无人车(UGV)结合,但该场景仍未得到充分探索:现有无训练方法仅能解决单智能体任务,未提供协同机制;近期CARLA-Air评估显示,5种最先进的VLA模型均未实现稳定的协同行为;朴素语义通信或双向耦合甚至会降低性能。本文提出AGC-VLN(空地协同VLN),是首个空地协同VLN的无训练基线。核心思路是无训练方法将导航分解为基于视觉语言模型(VLM)的语义推理与确定性几何执行,从而形成协同接口:无人机在全局视角下,将地面无人车报告的位姿及以VLM锚定的目标渲染为带有距离标签的CAR/GOAL标记,生成共享鸟瞰图。地面无人车从该图获取第一人称视角无法提供的全局空间上下文,用冻结的VLM规划沿道路的路径并执行闭环控制;同时无人机运行3D-SPF(SPF的空间搜索升级算法,用于在俯视视角下定位目标并飞向目标)。在CARLA-Air的Town10HD场景的100个闭环回合中,AGC-VLN的联合成功率达77.0%,相较于较弱的单智能体(无人机,50.0%)实现了27.0%的协同增益,且比最强的已发表单智能体基线(Travel UAV,53.0%)高出24.0个百分点,其性能源于无人机全局视角与地面无人车沿道路执行的互补性。项目页面:this https URL。
英文摘要
Air-ground collaborative Vision-and-Language Navigation (VLN) pairs an unmanned aerial vehicle (UAV) with a global bird's-eye view and an unmanned ground vehicle (UGV) with a local first-person view, yet the setting remains largely unexplored: existing training-free methods solve single-agent tasks but offer no collaboration mechanism, and a recent CARLA-Air evaluation found no stable cooperative behavior across five state-of-the-art VLA models; naive semantic communication or bidirectional coupling even degrades performance. We establish AGC-VLN (Air-Ground Collaborative VLN), the first training-free baseline for air-ground collaborative VLN. The key insight is that training-free methods decompose navigation into VLM-based semantic reasoning and deterministic geometric execution, exposing a collaboration interface: the UAV's global view, over which it renders the UGV's reported pose and the VLM-anchored target as CAR/GOAL markers with distance labels, yielding a shared bird's-eye map. From this map, the UGV acquires global spatial context its first-person view cannot provide, plans a road-following path with a frozen VLM, and executes it under closed-loop control; in parallel, the UAV runs 3D-SPF, a spatial-search upgrade of SPF that localizes the target in the downward view and flies toward it. On 100 closed-loop episodes in CARLA-Air's Town10HD scene, AGC-VLN reaches a 77.0% joint success rate, a collaboration gain of +27.0% over the weaker individual agent (the UAV, 50.0%), and exceeds the strongest published single-agent baseline (Travel UAV, 53.0%) by 24.0 points, stemming from the complementarity of the UAV's global view and the UGV's road-following execution. Project page: https://github.com/ZSN2024/AGC-VLN.
发表机构
- The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))
- Nanjing University of Posts and Telecommunications(南京邮电大学)
机构由 AI 辅助整理,请以论文原文为准。