arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2505.03460cs.RO

LogisticsVLN:基于智能体无人机的低空终端配送视觉-语言导航

LogisticsVLN: Vision-Language Navigation For Low-Altitude Terminal Delivery Based on Agentic UAVs

  • School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院)
  • State Key Laboratory of Multimodal Artificial Intelligence Systems, Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所多模态人工智能系统国家重点实验室)
  • Department of Engineering Science, Faculty of Innovation Engineering, Macau University of Science and Technology(澳门科技大学创新工程学院工程科学系)
  • China Ship Research and Development Academy(中国船舶科研 Academy)
  • Obuda University(奥布达大学)
  • State Key Laboratory for Management and Control of Complex Systems, Chinese Academy of Sciences(中国科学院复杂系统管理与控制国家重点实验室)

机构由 AI 辅助整理,请以论文原文为准。

Xinyuan Zhang, Yonglin Tian, Fei Lin, Yue Liu, Jing Ma, Kornélia Sára Szatmáry, Fei-Yue Wang

更新

AI总结:

针对无人机终端配送中现有VLN任务过于粗粒度的问题,本文提出基于多模态大语言模型的LogisticsVLN系统,并在CARLA中构建VLD数据集验证其可行性。

AI中文摘要:

智能物流需求的日益增长,尤其是细粒度终端配送,凸显了对基于自主无人机(UAV)配送系统的需求。然而,现有大多数最后一公里配送研究依赖地面机器人,而当前基于无人机的视觉-语言导航(VLN)任务主要关注粗粒度、长距离目标,使其不适合精确的终端配送。为弥合这一差距,我们提出LogisticsVLN,一个基于多模态大语言模型(MLLMs)构建的可扩展空中配送系统,用于自主终端配送。LogisticsVLN将轻量级大语言模型(LLMs)和视觉-语言模型(VLMs)集成到模块化流程中,用于请求理解、楼层定位、目标检测和动作决策。为支持这一新场景下的研究与评估,我们在CARLA模拟器中构建了视觉-语言配送(VLD)数据集。在VLD数据集上的实验结果表明了LogisticsVLN系统的可行性。此外,我们对系统各模块进行了子任务级评估,为提高基于基础模型的视觉-语言配送系统的鲁棒性和真实世界部署提供了有价值的见解。

英文摘要:

The growing demand for intelligent logistics, particularly fine-grained terminal delivery, underscores the need for autonomous UAV (Unmanned Aerial Vehicle)-based delivery systems. However, most existing last-mile delivery studies rely on ground robots, while current UAV-based Vision-Language Navigation (VLN) tasks primarily focus on coarse-grained, long-range goals, making them unsuitable for precise terminal delivery. To bridge this gap, we propose LogisticsVLN, a scalable aerial delivery system built on multimodal large language models (MLLMs) for autonomous terminal delivery. LogisticsVLN integrates lightweight Large Language Models (LLMs) and Visual-Language Models (VLMs) in a modular pipeline for request understanding, floor localization, object detection, and action-decision making. To support research and evaluation in this new setting, we construct the Vision-Language Delivery (VLD) dataset within the CARLA simulator. Experimental results on the VLD dataset showcase the feasibility of the LogisticsVLN system. In addition, we conduct subtask-level evaluations of each module of our system, offering valuable insights for improving the robustness and real-world deployment of foundation model-based vision-language delivery systems.

↑