发表机构
Saudi Data and Artificial Intelligence Authority (SDAIA)(沙特数据和人工智能管理局(SDAIA))
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对冻结双编码器在杂乱场景中检索稀有对象时全局嵌入不足的问题,提出无需训练的MINER框架,通过区域级嵌入和中心性校正重打分增强检索,并在ROCS基准上验证了其有效性。
AI 中文摘要
当查询命名杂乱场景中一个小的、视觉上从属的对象时,使用冻结双编码器的文本到图像检索性能会下降:单一的全局图像嵌入不足以代表局部视觉证据。我们提出了MINER,一个无需训练的推理框架,通过一小部分区域级嵌入和一种中心性校正的相似度重新评分来增强冻结双编码器的全局图像嵌入,从而恢复全局池化低估的视觉证据。为了评估这一设置,我们引入了ROCS,一个基于Flickr30K和MS COCO高杂乱子集构建的基准,其图像被重新标注以命名一个低显著性对象。在CLIP、SigLIP和SigLIP 2上的实验表明,MINER在每一个骨干网络上、在ROCS和标准分割上都提高了检索性能。分析表明,这些提升主要来自更广泛的空间覆盖而非精确的裁剪位置,揭示了一种从冻结表示中恢复局部证据的简单且通用的方法。代码:https://github.com/aalquwayfili/MINER。数据集:https://huggingface.co/datasets/aalquwayfili/ROCS。
英文摘要
Text-to-image retrieval with frozen dual encoders degrades when the query names a small, visually subordinate object in a cluttered scene: a single global image embedding underrepresents the localized visual evidence. We present MINER, a training-free inference framework that augments a frozen dual encoder's global image embedding with a small bank of region-level embeddings and a hubness-correcting similarity rescoring, recovering visual evidence that global pooling underweights. To evaluate this setting, we introduce ROCS, a benchmark built from high-clutter subsets of Flickr30K and MS COCO whose images are re-captioned to name a single low-salience object. Experiments on CLIP, SigLIP, and SigLIP 2 show that MINER improves retrieval on every backbone, on ROCS and on the standard splits. Analyses show that these gains come primarily from broader spatial coverage rather than precise crop placement, revealing a simple and general way to recover localized evidence from frozen representations. Code: https://github.com/aalquwayfili/MINER. Dataset: https://huggingface.co/datasets/aalquwayfili/ROCS.
CommentsAccepted at ACML 2026 (PMLR). 29 pages, 10 figures