arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

DroneGround:利用合成数据和接地视觉-语言模型进行开放词汇无人机载荷表征

DroneGround: Open-Vocabulary Drone Payload Characterization Using Synthetic Data and Grounded Vision-Language Models

Ami Pandat, Rajasekhar Punna, Gopika Vinod, Rohit Shukla

arXiv 2609.07780首次发表:更新:

发表机构

Homi Bhabha National Institute; Bhabha Atomic Research Centre(霍米·巴伯哈国家学院; 巴巴原子研究中心)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对无人机载荷表征在远距离和分布偏移下的挑战,提出基于合成数据和接地视觉-语言模型的两阶段开放词汇框架DroneGround,显著提升F1分数与泛化能力。

AI 中文摘要

自动化无人机监视对于公共安全、关键基础设施保护和受限空域监控已变得日益重要。尽管现有的基于视觉的系统在无人机检测和跟踪方面表现出色,但在远距离成像条件下,由于标注的真实世界数据集有限以及部署中遇到的显著分布偏移,可靠的载荷表征仍然极具挑战性。现有方法将载荷表征视为闭集目标检测问题,限制了其识别未见过的载荷并泛化到训练分布之外的能力。为应对这些挑战,我们使用Unreal Engine 5和Cosys-AirSim生成了逼真的合成无人机载荷数据集,并提出了DroneGround:接地视觉-语言载荷表征,这是一个用于稳健开放词汇载荷分析的两阶段框架。DroneGround首先采用YOLO26s检测器定位无人机并提取以无人机为中心的图像裁剪,随后由经LoRA微调的PaliGemma视觉-语言模型生成检测到的无人机及其附下载荷的语义描述,从而超越预定义类别实现开放词汇载荷表征。基于遮挡的接地模块通过识别负责生成描述的图像区域,进一步提供可解释的载荷定位。在合成和真实世界无人机图像上的广泛实验表明,DroneGround在合成到真实的分布偏移下显著提高了稳健性,将F1分数从82.5%提升至96.3%,优于传统的闭集载荷检测器,同时在对未见载荷类别的泛化方面表现更佳(F1为80.4%对42.7%)。数据集和代码将在论文被接收后发布。

英文摘要

Automated drone surveillance has become increasingly important for public safety, critical infrastructure protection,and restricted airspace monitoring. While existing vision-based systems achieve strong performance for drone detection and tracking, reliable payload characterization remains highly challenging under long-range imaging conditions due to limited availability of annotated real-world datasets, and substantial distribution shifts encountered during deployment. Existing approaches formulate payload characterization as a closed-set object detection problem, limiting their ability to recognize previously unseen payloads and generalize beyond the training distribution. To address these challenges, we generate a photorealistic synthetic drone-payload dataset using Unreal Engine 5 and Cosys-AirSim and propose DroneGround: Grounded Vision-Language Payload Characterization, a two-stage framework for robust open-vocabulary payload analysis. DroneGround first employs a YOLO26s detector to localize drones and extract drone-centric image crops, followed by a LoRA-fine-tuned PaliGemma vision-language model that generates seman- tic descriptions of the detected drones and their attached payloads, enabling open-vocabulary payload characterization beyond predefined categories. An occlusion-based grounding module further provides interpretable payload localization by identifying image regions responsible for the generated descriptions. Extensive experiments on both synthetic and real-world drone imagery demonstrate that DroneGround substantially improves robustness under synthetic-to-real distribution shifts, outperforming a conventional closed-set payload detector by improving the F1-score from 82.5% to 96.3%, while achieving significantly better generalization to previously unseen payload categories (80.4%versus 42.7% F1). Dataset and code will be released upon request.

CommentsAccepted at RVS-SE, British Machine Vision Conference, 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑