arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

面向野外无人机的主动目标检测:大规模数据集、基准和方法

Toward Active Object Detection for UAVs in the Wild: A Large-Scale Dataset, Benchmark and Method

Tianpeng Liu, Xinhua Jiang, Li Liu, Qinmu Shen, Siwei Tang, Zhen Liu, Yongxiang Liu

arXiv 2607.09078首次发表:更新:

AI 中文总结

针对无人机主动目标检测缺乏高质量数据集和基准的问题,提出ATRNet-LUDO数据集并建立评估基准。利用联合嵌入预测架构构建世界模型,提出AOD-JEPA,经实验验证其有效性和优越性,推动无人机-地面主动目标检测领域研究。

AI 中文摘要

目标检测是众多无人机应用中的基本组成部分,但长期受遮挡或目标像素稀缺等问题困扰。主动目标检测(AOD)提供了新范式,然而基于无人机的AOD研究因缺乏高质量数据集和基准而稀缺。本文提出ATRNet-LUDO,首个用于无人机-地面主动目标检测的大规模真实世界数据集,含多视图全景多目标航空图像等。基于此建立评估基准,多数现有AOD策略依赖深度强化学习但泛化性差。利用联合嵌入预测架构构建世界模型,提出AOD-JEPA并经实验验证其有效性和优越性,有望推动该领域研究。

英文摘要

Object detection is a fundamental component in numerous Unmanned Aerial Vehicle (UAV) applications, yet it has long been plagued by hindrances like occlusion or target pixel scarcity. Active Object Detection (AOD) provides a novel paradigm to address these challenges via active vision, while UAV-based AOD research remains scarce due to the lack of high-quality datasets and benchmarks for algorithm development and evaluation. To fill this gap, this paper presents ATRNet-LUDO, the first large-scale real-world dataset for UAV-Ground Active Object Detection (UGAOD). It contains 121,000 multi-view panoramic multi-target aerial images and 1.21 million local single-target slices, covering 10 vehicle targets across 40 scenarios. It enables the construction of diverse training and testing environments for UAV agent interaction and active observation policy learning. Based on this dataset, we establish a comprehensive evaluation benchmark for AOD policy learning methods. Most existing AOD policies rely on Deep Reinforcement Learning (DRL) but suffer from poor generalization. Evaluations on our benchmark reveal a significant generalization gap between training and testing performance, highlighting an urgent need for solutions. To this end, we leverage the Joint Embedding Predictive Architecture (JEPA) to construct a world model that enhances state representation learning, and propose AOD-JEPA by incorporating AOD-specific prior knowledge. Extensive experiments validate its effectiveness and superiority. We hope ATRNet-LUDO and the benchmark will advance research in the UGAOD field. The dataset and code are soon available at https://github.com/Leo000ooo/LUDO_dataset.

Comments18 pages, 19 figures, 5 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑