arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

RSFusionDet:水下RGB-声呐多模态目标检测

RSFusionDet: Underwater RGB-Sonar Multimodal Object Detection

Zhuoyan Liu, Yihan Wang, Bo Wang, Bing Wang, Ye Li

arXiv 2608.25367首次发表:更新:

发表机构

Harbin Engineering University; Hong Kong Polytechnic University(哈尔滨工程大学; 香港理工大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对水下单模态目标检测的传感器成像缺陷,本文构建RSFusion数据集并提出RSFusionDet模型,通过CAFusion模块与OMHead、OMLoss实现跨模态特征融合与目标匹配,在RSFusion数据集上性能优于基线模型DINO。

AI 中文摘要

水下单模态目标检测面临诸多传感器成像挑战,例如光学图像受水下噪声和可视距离限制,声呐图像则受目标结构信息较少的制约;而光学图像拥有丰富的目标结构信息,声呐图像受水下噪声影响更小且可视距离更长,光学(RGB模态)与声呐(Sonar模态)图像在水下场景中存在互补信息。本文构建了RGB-声呐多模态目标检测数据集RGB-Sonar Fusion(RSFusion),并为该基准数据集设计了评估指标;同时提出了RGB-Sonar Fusion Detector(RSFusionDet),为RGB-声呐多模态目标检测设计了全新的结果表达形式。本文分析了RGB和声呐模态信息的特征,设计了跨注意力融合(Cross-Attention Fusion,CAFusion)模块以融合RGB-声呐空间错位特征,还设计了目标匹配头(Object Matching Head,OMHead)与目标匹配损失(Loss,OMLoss)来匹配RGB-声呐模态中的相同目标。在RSFusion数据集上,RSFusionDet取得了76.4/48.6的目标检测AP(RGB/声呐)、83.4的目标匹配F1-Score,优于其他目标检测模型;与基线模型DINO相比,本方法在RGB/声呐AP上分别提升0.7/1.4,同时提供可靠的跨模态目标匹配。代码与数据集可通过该httpsURL公开获取。

英文摘要

Underwater unimodal object detection faces many challenges in sensor imaging, such as optical images limited by underwater noise and visible distance, and sonar images limited by less object structural information. While, optical images have rich object structural information, and sonar images are less affected by underwater noise and have a longer visible distance. Optical (RGB modality) and sonar (Sonar modality) images have complementary information underwater. In this paper, we create an RGB-Sonar multimodal object detection dataset, \textbf{R}GB-\textbf{S}onar \textbf{Fusion} (RSFusion) and propose evaluation metrics for the benchmark. And we propose the \textbf{R}GB-\textbf{S}onar \textbf{Fusion} \textbf{Det}ector (RSFusionDet) with a new RGB-Sonar multimodal object detection result expression for RGB-Sonar multimodal object detection. We analyze the features of RGB and Sonar modal information, and design a Cross-Attention Fusion (CAFusion) module to fuse RGB-Sonar spatial misalignment features and Object Matching Head (OMHead) with Loss (OMLoss) to match identical objects in RGB-Sonar modalities. Our RSFusionDet achieves 76.4/48.6 AP (RGB/Sonar) for object detection and 83.4 \(\text{F1-Score}_{match}\) for object matching, on RSFusion, which outperforms other object detection models. Compared with the DINO baseline, our method improves by 0.7/1.4 AP (RGB/Sonar) while simultaneously providing reliable cross-modal object matching. The code and datasets are publicly available at https://github.com/LEFTeyex/RSFusionDet.

Comments18 pages, 13 figures, 18 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑