arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

利用光学视觉基础模型进行SAR目标识别的跨模态学习

Cross-modal learning for SAR target recognition using optical vision foundation models

Lucas Hirsch, James R. Hopgood, Javid Khan, Yoann Altmann, Mike E. Davies

arXiv 2609.07753首次发表:更新:

发表机构

School of Engineering, University of Edinburgh; Leonardo UK; School of Engineering and Physical Sciences, Heriot-Watt University(爱丁堡大学工程学院; 英国莱昂纳多公司; 赫瑞瓦特大学工程与物理科学学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出跨模态EO到SAR原型对齐框架,利用冻结的DINOv3光学基础模型构建类别原型,指导SAR模型分类,在UNICORNv2数据集上显著提升准确率,证明光学基础模型可迁移至SAR识别。

AI 中文摘要

合成孔径雷达(SAR)因其多功能性、远距离和近乎全天候的工作能力,在广泛的成像应用中是一种重要的模态。然而,由于标记数据有限、SAR图像中强烈的斑点噪声以及SAR与更丰富的光学图像之间存在显著的域差距,自动目标识别(ATR)仍然是一个具有挑战性的问题。相比之下,电光(EO)图像受益于大规模数据集、更清晰的视觉结构和强大的基础模型。在这项工作中,我们研究了在光学数据上训练的视觉基础模型如何为SAR分类提供类别级监督。我们提出了一种跨模态EO到SAR原型对齐框架,其中基于DINOv3视觉基础模型的冻结EO编码器用于构建类别级光学原型,无需严格的EO/SAR配对。然后训练SAR模型对SAR图像进行分类,同时将其嵌入与相应的EO类别原型对齐。在推理时,SAR模型独立运行,无需访问光学图像。我们在UNICORNv2数据集上评估了我们的方法,该数据集是包含严重斑点噪声图像和严重类别不平衡的民用车辆EO和SAR数据集。EO原型对齐在SAR分类准确率上优于冻结DINOv3、仅SAR微调和未配对分布对齐基线,并且t-SNE可视化提供了训练后的SAR嵌入空间中类别间更清晰分离的定性证据。这些结果表明,尽管光学视觉基础模型是在可见光谱图像上训练的,但它们为SAR图像分类提供了可迁移的信息,为在具有挑战性的传感模态中使用大规模预训练视觉基础模型提供了一种实用方法。

英文摘要

Synthetic Aperture Radar (SAR) is an important modality in a wide range of imaging applications due to its versatile, long range and near all weather operating capabilities. However, Automatic Target Recognition (ATR) remains a challenging problem due to limited labelled data, the strong speckle in SAR images and the significant domain gap between SAR and more abundant optical imagery. In contrast, electro-optical (EO) imagery benefits from massive datasets, clearer visual structure and powerful foundation models. In this work, we investigate how vision foundation models trained on optical data can provide class level supervision for SAR classification. We propose a cross-modal EO to SAR prototype alignment framework in which a frozen EO encoder, based on a DINOv3 vision foundation model, is used to construct class level optical prototypes without requiring strict EO/SAR pairs. A SAR model is then trained to classify SAR images while aligning its embeddings to the corresponding EO class prototype. At inference time, the SAR model operates independently, without access to optical imagery. We evaluate our approach on the UNICORNv2 dataset, an EO and SAR dataset of civilian vehicles with heavily speckled images and severe class imbalance. EO prototype alignment improves SAR classification accuracy over frozen DINOv3, SAR only finetuning and unpaired distribution alignment baselines, and t-SNE visualizations provide qualitative evidence of clearer separation among classes in the trained SAR embedding space. These results suggest that optical vision foundation models, despite being trained on visible spectrum imagery, provide transferable information for SAR image classification, offering a practical method for using large scale pretrained vision foundation models across challenging sensing modalities.

CommentsAccepted for presentation at SPIE Sensors + Imaging 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑