arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.17917cs.CVcs.AI

自动目标检测与识别的开箱即用技术对比研究

Comparative Study of Out-of-the-Box Technology for Automatic Target Detection and Recognition

Alma M. Liezenga, Lotte Nijskens, Henrik R. Baumann, Stefan Becker, Simon Bensberg, Niccolò Camarlinghi, Håvard R. Eiring, Alexander W. Johnsgaard, Tanel Liiv, … 展开作者

Alma M. Liezenga, Lotte Nijskens, Henrik R. Baumann, Stefan Becker, Simon Bensberg, Niccolò Camarlinghi, Håvard R. Eiring, Alexander W. Johnsgaard, Tanel Liiv, Giuseppe Martino, Matteo Marturini, Matthias Rapp, Jan Erik van Woerden, Alexander Wolpert, Hugo J. Kuijf

首次发表
浏览论文内容

中文总结 AI 辅助

本文对比测试YOLO系列、DETR框架等模型在军事相关数据集上的性能,发现更大模型表现更优,DETR模型优于YOLO,域外微调可提升A2G性能,而域内训练仍是构建高性能ATD/R系统的关键。

中文摘要 AI 辅助

自动目标检测与识别(ATD/R)对军事决策支持和(半)自主行动至关重要。目标检测与人工智能(AI)领域的最新进展显著提升了ATD/R的潜在性能,但公开军事数据集的稀缺性限制了这些系统的应用。为此,本文探索利用公开可用模型和民用数据集在军事场景中实现合理性能。我们在新获取的军事相关数据集上对多款最先进模型进行基准测试,包括6个迭代版本的YOLO系列和2种变体的DETR框架,该数据集包含军用车辆及各类挑战性场景,如不同程度的遮挡和小目标。我们同时验证了各模型的开箱即用版本,以及在VisDrone数据集上微调后的版本,VisDrone数据集包含小目标、空对地(A2G)视角及相关类别,有望推广至我们的军事ATD/R任务。我们在空对地(A2G)和地对地(G2G)视角、目标尺寸、模型尺寸维度,使用mAP@0.5和mAP@0.5:0.95指标对比模型性能,以了解模型的实时能力。主要发现为:(1)更大的模型性能优于更小的模型;(2)基于DETR的模型相比YOLO系列表现出良好的结果;(3)在域外A2G数据集上微调模型可提升其A2G性能,并小幅提升小目标检测性能;(4)所有模型在A2G场景下检测小目标仍存在困难。我们得出结论,尽管目标检测领域取得了最新进展,但域内训练对于构建高性能ATD/R系统仍然至关重要。

英文摘要

Automatic Target Detection and Recognition (ATD/R) is critical for military decision support and (semi-)autonomous operations. Recent advances in object detection and artificial intelligence (AI) significantly boosted the potential performance of ATD/R. However, the scarcity of publicly available military datasets limits the application of these systems. As a solution, this paper explores the use of publicly available models and civilian datasets to achieve reasonable performance in military contexts. We benchmark several state-of-the-art models, including six iterations of the YOLO series and two variations on the DETR framework, on a newly acquired military relevant dataset. This dataset features military vehicles and challenging circumstances, including various degrees of occlusions and small targets. The out-of-the-box version of each model is validated alongside a version finetuned on the VisDrone dataset. This dataset features small objects, an Air-to-Ground (A2G) perspective and relevant classes, potentially generalizing to our military ATD/R task. We compare the performance of the models using mAP@0.5 and mAP@0.5:0.95, across A2G and Ground-to-Ground (G2G) perspective, target size and model size, giving insight into the real-time capabilities of models. Our main findings are: (1) bigger models outperform smaller models, (2) DETR-based models show promising results compared to the YOLO series,(3) fine-tuning models on an out-of-domain A2G dataset, improves their A2G performance and slightly improves their performance on small objects, but (4) all models still struggle with detecting small objects in an A2G scenario. We conclude that, despite recent advances in object detection, in-domain training is still crucial for creating capable ATD/R systems.

补充信息

↑