arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

帧级标签在热成像视频中小型无人机点检测中的能与不能

What Frame-Level Labels Can and Cannot Do for Small-UAV Point Detection in Thermal Video

Wonbin Son, Gyumum Choi, Junil Seo, Hyungjoon Kim

arXiv 2610.07705首次发表:更新:

AI 中文总结

本研究探讨利用帧级存在/不存在标签训练小型无人机热成像点检测器的能力与局限,分析训练阶段、标签分配等因素,为标签分配与虚警控制提供指导。

AI 中文摘要

无人机(UAV)的日益普及提高了基于图像的无人机检测的重要性。基于学习的检测器使用图像和标注进行训练,标注类型决定了训练期间可用的信息。我们专注于在传感器或场景变化使得额外训练的空间标注变得繁琐时,利用帧级目标存在/不存在标签来学习定位。我们分析了一种现有架构的检测能力、学习行为和潜在应用,该架构用于小型无人机的点检测,使用存在/不存在标签训练,且无需外部检测器。该架构冻结通过分类学习到的空间特征,并使用相同的帧标签训练一个读出器,以生成空间得分图和点检测。在两个热红外数据集CST Anti-UAV和Anti-UAV410上,我们评估了在虚警约束下的定位命中率和检测率,分析了训练阶段、标签分配、合成和模型配置的影响,并与边界框检测器进行了比较。我们还探索了在机载目标跟踪(AOT)上使用其可见光图像和帧标签的潜在应用。分类训练增强了与目标相关的空间响应,而读出器训练有助于一致地提取这些响应。在更多视频中分配相似的标签数量产生了更高的定位命中率,而合成效果因数据集和评估标准而异。较高的定位命中率并不总是改善虚警约束下的检测,并且在目标信号相对于背景变化较弱以及跨数据集迁移时仍存在失败。这些发现为标签分配、空间表示和读出器、合成以及虚警控制提供了指导。

英文摘要

The growing use of unmanned aerial vehicles (UAVs) has increased the importance of image-based UAV detection. Learning-based detectors are trained on imagery and annotations, with annotation type determining the information available during training. We focus on learning localization from frame-level target presence/absence labels when sensor or scene changes make spatial annotations for additional training burdensome. We analyze the detection capability, learning behavior, and potential applications of an existing architecture for point detection of small UAVs, trained with presence/absence labels and requiring no external detector. The architecture freezes spatial features learned through classification and trains a readout with the same frame labels to produce spatial score maps and point detections. On two thermal infrared datasets, CST Anti-UAV and Anti-UAV410, we evaluate localization hit rates and detection rates under false-alarm constraints, analyze the effects of training stages, label allocation, synthesis, and model configuration, and compare with bounding-box detectors. We also explore potential applications on Airborne Object Tracking (AOT) using its visible-light imagery and frame labels. Classification training strengthened target-related spatial responses, while readout training helped extract them consistently. Distributing similar label counts across more videos yielded higher localization hit rates, while synthesis effects varied by dataset and evaluation criterion. Higher localization hit rates did not always improve detection under false-alarm constraints, and failures remained when target signals were weak relative to background variation and under cross-dataset transfer. These findings provide guidance on label allocation, spatial representations and readouts, synthesis, and false-alarm control.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑