arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

共显着目标检测的秩一致集合推理

Rank-Consistent Set Reasoning for Co-Salient Object Detection

Yuan Xiang, Matteo Rossi, Yingzhou Chen

arXiv 2609.13706首次发表:更新:

发表机构

University of California, Los Angeles (UCLA); Polytechnic University of Turin(加州大学洛杉矶分校; 都灵理工大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对共显着目标检测,提出秩一致集合推理(RCSR)框架,将图像组建模为无序集合并通过排序与截尾统计聚合,抑制干扰并提升鲁棒性,在多个基准上验证了有效性。

AI 中文摘要

共显着目标检测(Co-SOD)要求模型找到在单个图像中显着且受图像组支持的感兴趣区域。我们提出了秩一致集合推理(RCSR),一种有监督的密集预测框架,该框架将图像组建模为无序集合,而非图像序列或语义标签。其核心思想是在每个图像尺度上,对每个空间区域与少量学习到的组槽(group slots)的一致性程度进行排序,并使用稳健的截尾统计量聚合这些排序。这抑制了偶然的成对匹配,并防止一个非典型的组成员主导共享表示。集合编码器直接从多尺度视觉特征构建组槽,而秩一致性门控则衡量候选区域的排序在组成员之间是否稳定。门控槽与逐图像特征联合解码以生成共显着图。该模型不包含自然语言分支、开放词汇检测器或外部分割模型。我们进一步引入了组排列目标函数和困难干扰物增强,使模型学习集合级目标的属性,而非记忆图像顺序或孤立的视觉显着性。我们为CoCA、CoSal2015和CoSOD3k制定了一套评估协议,并进行了组大小鲁棒性、干扰物拒绝、顺序不变性和跨数据集迁移的测试。

英文摘要

Co-salient object detection (Co-SOD) requires a model to find foreground regions that are salient in individual images and supported by the image group. We present \emph{Rank-Consistent Set Reasoning} (RCSR), a supervised dense-prediction framework that models a group as an unordered set rather than as a sequence of images or a semantic label. The core idea is to rank how strongly each spatial region agrees with a small collection of learned group slots at every image scale, and to aggregate these ranks with a robust trimmed statistic. This suppresses accidental pairwise matches and prevents one atypical group member from dominating the shared representation. A set encoder builds group slots directly from multi-scale visual features, while a rank-consistency gate measures whether the ordering of candidate regions is stable across group members. The gated slots are decoded jointly with per-image features to produce co-saliency maps. The model contains no natural-language branch, no open-vocabulary detector, and no external segmentation model. We further introduce a group permutation objective and hard-distractor augmentation so that the model learns the properties of a set-level target rather than memorizing image order or isolated visual saliency. We formulate an evaluation protocol for CoCA, CoSal2015, and CoSOD3k, together with tests of group-size robustness, distractor rejection, order invariance, and cross-dataset transfer.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑