SonarLLM:一种用于水下感知的原生声呐-光学多模态大语言模型
SonarLLM: A Native Sonar--Optical Multimodal Large Language Model for Underwater Perception
浏览论文内容
中文总结 AI 辅助
针对现有多模态大语言模型不适用于水下声呐-光学感知的问题,提出SonarLLM模型,结合专用编码器与分层融合机制,在SonarBench基准上实现了优异的水下感知性能,凸显了声呐的互补价值。
中文摘要 AI 辅助
可靠的水下感知需要在能见度变化的环境下利用互补传感技术:光学相机能捕捉外观和语义信息,但在水体浑浊时性能会迅速下降;而成像声呐可保留几何结构,却具有独特的距离-方位结构和声学伪影。现有多模态大语言模型(MLLM)主要基于光学编码器构建,因此不适用于建模声呐或自适应利用声呐-光学的互补性。我们提出SonarLLM,一种将声呐视为原生感知模态的声呐-光学多模态大语言模型,它结合了声呐专用编码器、模态专用的物理感知特征增强模块以及可靠性感知的分层融合机制,以对齐声学结构与光学语义,并根据传感质量的变化动态调整两者的贡献。我们还推出了SonarBench,这是一个配对基准数据集,涵盖识别、计数、视觉问答和字幕生成四项任务,以及仅声呐、仅光学、融合三种输入设置。SonarBench通过固定场景和声呐观测,同时改变光学退化程度,实现了跨模态互补性的可控测量。SonarLLM在仅声呐的识别、计数和视觉问答任务中达到72.0%的宏观准确率,比最强基线高出34.4个百分点;在融合输入设置下准确率为68.7%,比最佳基线高出25.1个百分点。在识别和计数任务中,随着浑浊度增加,融合输入相对于仅光学输入的增益从6.0个百分点增长至36.0个百分点,表明在光学退化受控制的情况下,声呐的互补价值不断提升。这些结果共同表明,稳健的异构感知不仅取决于添加声呐,还取决于根据声呐的传感特性对其进行表征和加权。
英文摘要
Reliable underwater perception requires complementary sensing under variable visibility. Optical cameras capture appearance and semantics but degrade rapidly with turbidity, whereas imaging sonar preserves geometry while exhibiting distinct range-azimuth structure and acoustic artifacts. Existing MLLMs, built primarily on optical encoders, are therefore ill-suited to model sonar or adaptively exploit sonar-optical complementarity. We propose SonarLLM, a sonar-optical MLLM that treats sonar as a native perceptual modality. It combines a sonar-specific encoder, modality-specific physics-aware feature enhancement, and reliability-aware hierarchical fusion to align acoustic structure with optical semantics and dynamically adjust their contributions as sensing quality changes. We also introduce SonarBench, a paired benchmark that spans four tasks: recognition, counting, visual question answering, and captioning; and, across the benchmark, three input settings: sonar-only, optical-only, and fusion. By fixing the scene and sonar observation while varying optical degradation, SonarBench enables controlled measurement of cross-modal complementarity. SonarLLM achieves 72.0% macro accuracy across sonar-only recognition, counting, and VQA, outperforming the strongest baseline by 34.4 percentage points, and 68.7% under fusion, exceeding the best baseline by 25.1 points. For recognition and counting, the fusion-over-optical gain grows from 6.0 to 36.0 points as turbidity increases, indicating the increasing complementary value of sonar under controlled optical degradation. Together, these results show that robust heterogeneous perception depends not only on adding sonar, but on representing and weighting it according to its sensing characteristics.
发表机构
- Faculty of Information Engineering and Automation, Kunming University of Science and Technology(昆明理工大学信息工程与自动化学院)
- Yunnan Key Laboratory of Artificial Intelligence(云南省人工智能重点实验室)
机构由 AI 辅助整理,请以论文原文为准。