AI 中文总结
该研究针对MLLMs时代SOD的能力不匹配问题,提出无训练框架FOCUS,在13类SOD基准上超越SOTA方法,推动SOD从特定任务监督转向零样本前景组织。
AI 中文摘要
多模态大语言模型(MLLMs)的零样本能力正推动显著性目标检测(SOD)突破特定任务监督的局限。为了让MLLMs脱离传统基于掩码的评估框架,我们将SOD拆解为定位与分割任务,并用语义短语、边界框和属性重新构建数据集,建立了针对MLLM显著性感知的诊断基准SaliLLM。SaliLLM揭示了显著的能力不匹配:MLLMs在定位任务上优于当前最优(SOTA)方法,但在分割任务上仍明显更弱。进一步分析将该差距主要归因于MLLMs与标注在前景数量、粒度和范围上的不匹配。基于这一诊断,我们将零样本SOD重新定义为协议对齐的前景组织,并提出首个无训练框架FOCUS,该框架受格式塔原理启发,采用协作注意力实现统一SOD。FOCUS结合了对协议条件前景粒度的自上而下贝叶斯惊喜校准,以及对自监督特征诱导的以实体为中心感知流形上的MLLMs证据的自下而上传播,生成连贯的目标范围作为通用分割器的提示。在13个RGB、RGB-D和RGB-T SOD基准上,FOCUS无需训练即可普遍超越SOTA方法,与全监督、弱监督和自监督方法相比,平均绝对误差分别降低了11%、34%和48%。我们的发现表明SOD正迎来复兴:从特定任务监督转向零样本前景组织。代码见补充材料。
英文摘要
The zero-shot capabilities of multimodal large language models (MLLMs) are pushing salient object detection (SOD) beyond task-specific supervision. To disentangle MLLMs beyond conventional mask-based evaluation, we decompose SOD into localization and segmentation, and re-engineer datasets with phrases, boxes, and attributes, establishing a diagnostic benchmark for MLLM saliency perception (SaliLLM). SaliLLM uncovers a striking capability mismatch: MLLMs outperform state-of-the-art (SOTA) methods in localization, yet remain substantially weaker in segmentation. Further analyses attribute this gap primarily to mismatches between MLLMs and annotations over foreground cardinality, granularity, and extent. Motivated by this diagnosis, we recast zero-shot SOD as protocol-aligned Foreground Organization and introduce the first training-free framework that leverages Gestalt-inspired Collaborative attention for Unified SOD (FOCUS). FOCUS couples top-down Bayesian-surprise calibration of protocol-conditioned foreground granularity with bottom-up propagation of MLLMs evidence over entity-centric perceptual manifolds induced by self-supervised features, yielding coherent object extents as prompts for a general segmenter. Across 13 RGB, RGB-D, and RGB-T SOD benchmarks, FOCUS generally surpasses SOTA methods without training, reducing mean absolute error by 11\%, 34\%, and 48\% compared with fully, weakly, and self-supervised methods, respectively. Our findings signal the renaissance of SOD: from task-specific supervision to zero-shot foreground organization. Code is available in the supplementary material.
Comments10 pages, 4 figures, conference