退化视觉条件下水下机器人的鲁棒跨模态基础模型感知
Robust Cross-Modal Foundation Model Perception for Underwater Robots under Degraded Visual Conditions
浏览论文内容
中文总结 AI 辅助
该研究针对水下机器人视觉退化问题,提出感知退化的视觉-声纳门控融合方法,在极端退化下使平衡准确率较 DINOv2 基线提升 33.5%,实现鲁棒跨模态感知。
中文摘要 AI 辅助
可靠的水下机器人感知仍存在困难,因为光学图像会因浑浊度、波长相关衰减、低光照、散射和模糊而退化。尽管声纳提供了受光学能见度影响较小的互补信息,但现有的视觉-声纳研究大多集中在特征对齐和标称检测性能上。我们研究了视觉可靠性恶化时的跨模态鲁棒性,并评估在严重退化情况下,预训练的视觉基础模型表示是否可以通过声纳进行补充。我们使用冻结的 DINOv2 作为视觉编码器,构建了从干净到极端视觉条件的受控五级基准。我们比较了传统视觉检测、冻结基础模型表示、声纳上下文、固定多模态融合、在干净数据上训练的自适应门控,以及感知退化的门控融合。我们的方法在整个退化范围内训练融合机制,同时保持视觉和声纳编码器冻结,从而允许模态贡献进行自适应调整,而无需微调预训练的骨干网络。在极端组合退化下,DINOv2 基线的平衡准确率为 0.4610,而感知退化的视觉-声纳融合达到 0.6152,相对提升了 33.5%。学习到的声纳贡献从干净条件下的 14.2% 增加到极端退化下的 41.3%,证明了跨模态依赖的自适应重新分配。融合在严重浑浊和模糊下提供最大增益,而仅颜色衰减几乎没有额外益处。这些结果表明,基础模型表示在严重信息丢失下仍然有价值但不足,而显式使融合适应模态可靠性可改善鲁棒的水下多模态感知。
英文摘要
Reliable underwater robotic perception remains difficult because optical imagery degrades under turbidity, wavelength-dependent attenuation, low illumination, scattering, and blur. Although sonar provides complementary information that is less affected by optical visibility, prior visual-sonar research has largely focused on feature alignment and nominal detection performance. We investigate cross-modal robustness as visual reliability deteriorates and assess whether pretrained visual foundation-model representations can be complemented by sonar under severe degradation. We use frozen DINOv2 as the visual encoder and construct a controlled five-level benchmark ranging from clean to extreme visual conditions. We compare conventional visual detection, frozen foundation-model representations, sonar context, fixed multimodal fusion, clean-trained adaptive gating, and degradation-aware gated fusion. Our method trains the fusion mechanism across the full range of degradation while keeping the visual and sonar encoders frozen, allowing modality contributions to adapt without fine-tuning the pretrained backbone. Under extreme combined degradation, the DINOv2 baseline achieves 0.4610 balanced accuracy, while degradation-aware visual-sonar fusion reaches 0.6152, a 33.5% relative improvement. The learned sonar contribution increases from 14.2% under clean conditions to 41.3% under extreme degradation, demonstrating adaptive redistribution of cross-modal reliance. Fusion provides the largest gains under severe turbidity and blur, whereas color attenuation alone yields little additional benefit. These results show that foundation-model representations remain valuable but insufficient under severe information loss, and that explicitly adapting fusion to modality reliability can improve robust underwater multimodal perception.
发表机构
- College of Science and Technology, North Carolina A & T State University(北卡罗来纳农工州立大学科学技术学院)
机构由 AI 辅助整理,请以论文原文为准。