基于大型深度神经网络改进合成孔径声呐图像的自动目标识别
Improved Automatic Target Recognition in Synthetic Aperture Sonar Imagery Using Large Deep Neural Networks
浏览论文内容
中文总结 AI 辅助
该研究对比CNN与Transformer等DNN架构的性能,探究网络规模、预训练等因素对合成孔径声呐自动目标识别的影响,旨在生成高性能模型并提供训练路线图。
中文摘要 AI 辅助
合成孔径声呐(SAS)中的自动目标识别(ATR)任务主要由深度神经网络(DNN)主导。大多数SAS-ATR模型使用卷积神经网络(CNN)架构,而基于Transformer的架构在文献中出现较少,尽管Transformer在通用计算机视觉(CV)研究中处于领先地位。此外,研究人员在尝试通过数据增强、使用多种成像模态的预训练权重等方法克服标记训练数据稀缺带来的挑战时,结果参差不齐。本研究中,我们比较了现代CNN和基于Transformer的DNN的性能,以确定哪种架构和训练配置能在SAS-ATR中产生最高性能。我们研究了网络规模、架构、预训练方法、数据增强及其他正则化形式对SAS-ATR性能的影响,重点是生成性能最高的模型,并为训练最先进的SAS-ATR模型提供路线图。
英文摘要
Automatic Target Recognition (ATR) in Synthetic Aperture Sonar (SAS) is a task largely dominated by deep neural networks (DNNs). Most SAS-ATR models use convolutional neural network (CNN) architectures whereas transformer-based architectures have had much less representation in the literature despite being state of the art in general computer vision (CV) research. Additionally, researchers have had mixed results in attempting to overcome challenges presented by a scarcity of labeled training data by using methods such as data augmentation and the use of pretrained weights from a variety of imaging modalities. In this work, we compare the performance of modern CNN and transformer-based DNNs to determine which architecture and training configurations elicit the highest performance in SAS-ATR. We investigate how network size, architecture, pretraining method, data augmentation and other forms of regularization affect SAS-ATR performance with a focus on producing the highest-performing model and providing a roadmap for training state-of-the-art SAS-ATR models.