arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

通过基于 SAM 的伪标注和基于状态空间交互的扩散进行弱监督 RGB-D 显著目标检测

Weakly-Supervised RGB-D Salient Object Detection via SAM-driven Pseudo Annotation and State Space Interaction-based Diffusion

Wenqi Si, Gongyang Li, Shixiang Shi, Weisi Lin

arXiv 2607.15041首次发表:更新:

发表机构

School of Communication and Information Engineering, Shanghai University; School of Computer Science and Engineering, Nanyang Technological University(上海大学通信与信息工程学院; 南洋理工大学计算机科学与工程学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究弱监督 RGB-D 显著目标检测,提出由 SAM-PAG 和 $S^2$Diff 组成的方法,前者扩展稀疏涂鸦为伪标注,后者在条件信息引导下细化显著图,二者合作使方法性能优异。

AI 中文摘要

弱监督 RGB-D 显著目标检测旨在减轻像素级标注的沉重负担。但涂鸦标注缺乏物体的结构和细节,导致显著图不准确。本文提出一种新的涂鸦监督 RGB-D 显著目标检测方法,由基于分割一切模型(SAM)的伪标注生成方法(SAM-PAG)和基于状态空间交互的条件扩散模型($S^2$Diff)组成。SAM-PAG 利用先进的 SAM 通过双分支结构和分割掩码一致性将稀疏涂鸦扩展为密集像素级伪标注。$S^2$Diff 采用扩散模型在条件信息引导下迭代细化有噪声的显著图以生成准确显著图。其核心在于获取条件特征和显著图去噪。通过 SAM-PAG 和 $S^2$Diff 的紧密合作,该方法在七个数据集上优于相关涂鸦监督方法,与全监督方法相比也具有竞争力。

英文摘要

Weakly-supervised RGB-D Salient Object Detection (SOD) is explored to reduce the heavy burden of pixel-level annotations. But scribble annotations lack the structure and details of objects, resulting in inaccurate saliency maps. In this paper, we propose a novel scribble-supervised RGB-D SOD method, consisting of a Segment Anything Model (SAM)-driven pseudo annotation generation method (\emph{SAM-PAG}) and a state space interaction-based conditional diffusion model (\emph{$S^2$Diff}). Specifically, SAM-PAG is tailored to address the issue of sparse supervision information. In SAM-PAG, we adopt the advanced SAM to expand sparse scribbles to dense pixel-level pseudo annotations through the dual-branch structure and the consistency of segmentation masks. In $S^2$Diff, we adopt the diffusion model to iteratively refine the noisy saliency maps with the guidance of conditional information, generating accurate saliency maps. Naturally, the core of our $S^2$Diff lies in the acquisition of conditional features and the denoising of saliency maps. For the former, we employ a cross-modal conditional generation module to interweave cross-modal features through frequency integration and implicit-explicit state space interaction, effectively achieving global conditional features. For the latter, we employ a context injection module to mitigate noise interference and to enhance object information with the conditional context. With the close cooperation of SAM-PAG and $S^2$Diff, our method outperforms relevant scribble-supervised methods and achieves competitive performance compared to fully-supervised methods on seven datasets. The code and results of our method are available at https://github.com/Switch457/WeakS2Diff_SOD.

Comments13 pages, 9 figures, accepted by IEEE TMM

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑