arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

S³AM:一种用于多模态显著目标检测的、具备可靠性校准频率适配器的单流SAM

S$^3$AM: A Single-Stream SAM with Reliability-Calibrated Frequency Adapter for Multi-modal Salient Object Detection

Ruichao Hou, Boyue Xu, Tongwei Ren, Dongming Zhou, Gangshan Wu, Jinde Cao

arXiv 2608.17475首次发表:更新:

发表机构

China Pharmaceutical University; Nanjing University; Yunnan University; Southeast University; Purple Mountain Laboratories(中国药科大学; 南京大学; 云南大学; 东南大学; 紫金山实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对多模态显著目标检测的冗余计算问题,提出融合可靠性校准频率适配器的单流SAM框架,仅用12.20M参数即实现竞争力性能,相关代码将公开。

AI 中文摘要

视觉基础模型近期通过参数高效微调与提示学习推动了多模态显著目标检测(MSOD)的发展。然而,现有适配Segment Anything Model(SAM)的MSOD方法常依赖双流编码器或辅助提示生成器,导致计算冗余。尽管单流方案可降低该成本,但早期融合可能会将含噪或未对齐的辅助高频线索传递至骨干网络。本文提出一种新颖的单流框架,将可靠性校准频率适配融入所采用的SAM骨干以实现MSOD,避免重复的基础骨干,同时显式控制辅助频率注入。具体而言,我们设计了混合频率专家模块,利用平稳小波变换分解各模态并聚合跨模态频率信息;进一步引入具备双门校准机制的可靠性校准频率适配器,在Transformer层级选择性传递校准后的残差,同时联合控制其注入强度与跨模态可靠性;超网络引导的语义-结构解码器则结合所采用骨干的语义掩码特征与基于Mamba的结构细节恢复。在RGB-D、RGB-T、RGB-NIR显著目标检测基准上的综合实验验证,该框架仅需12.20M可训练参数(占总参数的5.4%)即可实现具有竞争力的性能,代码将发布于指定链接。

英文摘要

Vision foundation models have recently advanced multi-modal salient object detection (MSOD) through parameter-efficient tuning and prompt learning. However, existing Segment Anything Model (SAM)-adapted MSOD methods often rely on dual-stream encoders or auxiliary prompt generators, leading to redundant computation. Although a single-stream alternative can reduce this cost, early fusion may also propagate noisy or misaligned auxiliary high-frequency cues through the backbone. In this paper, we propose a novel single-stream framework that integrates reliability-calibrated frequency adaptation into the adopted SAM backbone for MSOD. It avoids duplicated foundation backbones while explicitly controlling auxiliary frequency injection. Specifically, we design a mixture of frequency experts module, which uses the stationary wavelet transform to decompose each modality and aggregate cross-modal frequency information. We further introduce a reliability-calibrated frequency adapter with a dual-gate calibration mechanism, which selectively propagates the calibrated residual across transformer stages while jointly controlling its injection strength and cross-modal reliability. A hypernetwork-guided semantic-structural decoder then combines semantic mask features from the adopted backbone with Mamba-based structural detail recovery. Comprehensive experiments on RGB-D, RGB-T, and RGB-NIR salient object detection benchmarks validate that the proposed framework achieves competitive performance with only 12.20M trainable parameters, accounting for 5.4\% of the total parameters. The code will be available at https://github.com/xuboyue1999/SSSAM.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑