arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

MIMONet:用于显著目标检测的多尺度输入与多尺度输出网络

MIMONet: Multi-scale Input and Multi-scale Output Network for Salient Object Detection

Zhaojian Yao, Wei Gao, Tiesong Zhao, Hui Yuan, Sam Kwong

arXiv 2608.25733首次发表:更新:

发表机构

School of Electronic and Computer Engineering, Shenzhen Graduate School, Peking University; Peng Cheng Laboratory; College of Physics and Information Engineering, Fuzhou University; School of Control Science and Engineering, Shandong University; Department of Computer Science, City University of Hong Kong(北京大学深圳研究生院电子与计算机工程学院; 鹏城实验室; 福州大学物理与信息工程学院; 山东大学控制科学与工程学院; 香港城市大学计算机科学系)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对现有显著目标检测模型难以学习目标尺寸变化的问题,提出MIMONet,通过多尺度输入分支、MSP模块与JSL损失提升检测性能,在多数据集上取得更优结果,代码将公开。

AI 中文摘要

现有显著目标检测方法聚焦于多级特征的应用,旨在利用高层与低层特征的各自优势。然而,由于这些模型的输入为单尺度图像,其多级特征难以学习显著目标的尺寸变化知识。目标尺度变化学习对多尺度目标检测具有巨大潜力,但现有方法尚未充分探索这一点。为提升模型对不同尺寸目标的识别能力,受图像金字塔启发,我们提出了多尺度输入与多尺度输出网络(MIMONet)。在MIMONet中,我们为三张不同分辨率的图像提取多级特征,形成三个编码器分支,分支间会进行信息交换,该方法的优势在于,一个分支的特征可从另外两个分支的特征中学习目标尺寸变化的知识。此外,我们设计了多尺度感知(MSP)模块,该模块将输入特征层划分为多个不同分辨率的子层,在这些子层中捕获目标的多级结构信息,可使目标得到更充分的感知。针对网络训练,我们提出了联合显著损失(JSL),该损失可约束网络输出的多个显著图识别同一前景目标,并促使其边界得到清晰保留。实验结果表明,与现有模型相比,MIMONet在多个数据集上具备更强的检测能力,取得了更优的评估分数,我们的模型代码将予以公开。

英文摘要

The existing methods for saliency detection task focus on the application of multi-level features, aiming to take advantage of the respective strengths of high- and low-level features. However, because the inputs of these models are single-size images, their multi-level features have difficulty in learning the knowledge of size variations of salient objects. Object-scale variation learning has great potential for detecting multi-scale objects, which has not been fully explored by existing methods. To improve the recognition ability of a model for objects with different sizes, we are inspired by the image pyramid to propose a Multi-scale Input and Multi-scale Output Network (MIMONet). In MIMONet, we extract multi-level features for three images with different resolutions to form three encoder branches, and information will be exchanged between the branches. The advantage of this approach is that the features of one branch can learn the knowledge of target size variation from the features of the other two branches. In addition, we design a Multi-scale Perception (MSP) module, in which the input feature layer is divided into several sub-layers with different resolutions. Capturing the multi-level structure information of the objects in these sub-layers can make the objects more fully perceived. For network training, we propose a Joint Saliency Loss (JSL), which can constrain multiple saliency maps output by the network to identify the same foreground objects, and induce their boundaries to be preserved clearly. Experimental results show that MIMONet has stronger detection capabilities and harvests better evaluation scores on multiple datasets compared to existing models. The code of our model will be released.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑