BenthicDINO:用于视角不变侧扫声呐表征的物理感知自蒸馏方法
BenthicDINO: Physics-Informed Self-Distillation for View-Invariant Side-Scan Sonar Representations
浏览论文内容
中文总结 AI 辅助
本文提出BenthicDINO框架,基于DINOv3架构结合物理感知自蒸馏,通过物理增强与HSIC惩罚实现侧扫声呐表征的视角不变性,在S3Seg数据集上仅用10%标注数据即可达峰值性能的96%,取得71.4% mIoU与86.5%总体准确率。
中文摘要 AI 辅助
侧扫声呐(SSS)图像的自动感知受物理声学伪影严重阻碍,导致表征将海底固有反射率与瞬时观测几何结构不可分割地混合在一起。现有自监督学习(SSL)框架依赖为自然图像设计的数据增强,未考虑声学退化问题,也未明确强制视角不变性。为解决这一缺口,我们引入基于DINOv3架构的物理感知自蒸馏框架,采用ConvNeXt-v2-Tiny骨干网络以最大化数据效率。该方法通过两种核心机制实现视角不变性:一是模拟散斑噪声、距离相关衰减和辐射度校准误差的物理驱动增强;二是希尔伯特-施密特独立性准则(HSIC)惩罚项,明确将学习到的密集块特征与物理观测参数解耦。此外,我们提出跨网络四个阶段的密集分层特征融合策略,以在保留深层语义抽象的同时保留细粒度沉积物细节。大量评估表明,该框架无需人工标注即可将复杂海底地形自然聚类为稳定、无噪声的语义簇。在S3Seg数据集的监督下游任务中,融合后的表征展现出卓越的数据效率:仅使用10%的可用标注数据即可达到其绝对峰值性能的96%,最终实现71.4%的平均交并比(mIoU)和86.5%的总体准确率。
英文摘要
Automated perception in side-scan sonar (SSS) imagery is severely hindered by physical acoustic artifacts, resulting in representations that inextricably mix intrinsic seabed reflectivity with transient viewing geometries. Existing self-supervised learning (SSL) frameworks rely on augmentations designed for natural images, failing to account for acoustic degradation and explicitly enforce view-invariance. To address this gap, we introduce a physics-informed self-distillation framework built upon the DINOv3 architecture utilizing a ConvNeXt-v2-Tiny backbone to maximize data efficiency. The proposed methodology enforces view-invariance through two primary mechanisms: physically motivated augmentations that simulate speckle noise, range-dependent attenuation, and radiometric miscalibration; and a Hilbert-Schmidt Independence Criterion (HSIC) penalty that explicitly decouples learned dense patch features from physical viewing parameters. Furthermore, we propose a dense, hierarchical feature fusion strategy across all four network stages to preserve fine-grained sediment details alongside deep semantic abstractions. Extensive evaluation demonstrates that the framework natively groups complex benthic topographies into stable, noise-free semantic clusters without relying on manual annotations. During supervised downstream tasks on the S3Seg dataset, the fused representations exhibited exceptional data efficiency, achieving 96% of its absolute peak performance using only 10% of the available annotated data, ultimately reaching a mean Intersection over Union (mIoU) of 71.4% and an overall accuracy of 86.5%.
发表机构
- Computer Vision and Robotics Research Institute (ViCOROB)(计算机视觉与机器人研究学院(ViCOROB))
- University of Girona(赫罗纳大学)
机构由 AI 辅助整理,请以论文原文为准。