LoRA增强的对比学习与SAS视觉变换器
LoRA Enhanced Contrastive Learning with SAS Vision Transformers
浏览论文内容
中文总结 AI 辅助
本文提出三阶段参数高效框架,将DINOv3 ViT适配到SAS水下目标识别,证明仅LoRA适配即可显著提升AUPRC至0.679,而难负样本挖掘和SupCon细化无额外增益。
中文摘要 AI 辅助
利用合成孔径声纳(SAS)进行自动目标识别(ATR)可支持先进的海军能力,但深度学习受到目标图像稀缺、背景杂波和人在回路评估的限制。我们采用三阶段参数高效框架,将DINOv3视觉变换器(ViT)模型适配到水下SAS ATR。第一阶段使用低秩适配(LoRA)同时冻结ViT主干,弥合自然图像预训练与水下声传播之间的差距。第二阶段使用难负样本挖掘来强化决策边界,以对抗声学模仿物,包括类似人造目标的岩石和沉积构造。第三阶段使用监督对比学习(SupCon)来分离目标和杂波表示。我们使用任务级地理划分评估海上SAS数据,在85%测试召回率下比较所有分支,并重复每次比较三次随机种子。LoRA是主要效应,使用相同冻结主干将精确率-召回率曲线下面积(AUPRC)从0.300提升至0.679 ± 0.027。秩4在仅训练0.26%权重的情况下实现了这一结果。两个细化阶段均未超过其匹配对照:难负样本挖掘相对于等大小随机课程将AUPRC改变-0.0045 ± 0.0119,而SupCon相对于前一阶段将AUPRC改变+0.0002 ± 0.0096。这些零结果表明显式挖掘发生在编码器已拟合的数据上,且监督阶段已施加了大部分目标-杂波几何结构。一个高效的适配阶段就足够了;堆叠细化则不然。
英文摘要
Automatic target recognition (ATR) with synthetic aperture sonar (SAS) supports advanced naval capabilities, but deep learning is constrained by scarce target imagery, background clutter, and human-in-the-loop assessment. We adapt DINOv3 Vision Transformer (ViT) models to underwater SAS ATR using a three-stage parameter-efficient framework. Stage 1 uses Low-Rank Adaptation (LoRA) while freezing the ViT backbone, bridging the gap between natural-image pretraining and underwater acoustic propagation. Stage 2 uses hard-negative mining to strengthen the decision boundary against acoustic mimics, including rocks and sediment formations resembling man-made targets. Stage 3 uses Supervised Contrastive Learning (SupCon) to separate target and clutter representations. We evaluate at-sea SAS data using a mission-level geographic split, compare all arms at 85 percent test recall, and repeat each comparison over three random seeds. LoRA accounts for the primary effect, increasing area under the precision-recall curve (AUPRC) from 0.300 to 0.679 +/- 0.027 using the same frozen backbone. Rank 4 achieves this result while training only 0.26 percent of weights. Neither refinement stage exceeds its matched control: hard-negative mining changes AUPRC by -0.0045 +/- 0.0119 versus an equal-size random curriculum, and SupCon changes AUPRC by +0.0002 +/- 0.0096 versus the preceding stage. These null results indicate that mining occurred on data the encoder had already fit and that supervised stages had already imposed most target-clutter geometry. One efficient adaptation stage is sufficient; stacked refinement is not.
发表机构
- Florida Atlantic University(佛罗里达大西洋大学)
- Naval Surface Warfare Center Panama City Division(海军水面作战中心巴拿马城分部)
- Florida State University(佛罗里达州立大学)
机构由 AI 辅助整理,请以论文原文为准。