重新思考特定领域的基于文本的图像检索
Rethinking Text-Based Image Retrieval in Specific Domain
AI总结:
针对现有TBIR基准的单匹配假设无法适配监控等特定领域的问题,构建了SecMM-TBIR基准,提出SAFT框架解决特定领域TBIR的语义压缩问题,在SecMM-TBIR上平均提升mAP@20达7.8点且优化通用领域性能。
AI中文摘要:
视觉-语言表示学习的快速推进推动了基于文本的图像检索(TBIR)取得显著进展。然而,现有基准大多基于查询与图像间的唯一匹配假设构建,该假设在通用场景中有效,但无法反映特定领域(如监控)的实际系统性能——在这类领域,单个查询通常对应多个相关候选图像。为解决此局限,我们设计了特定领域多匹配基于文本的图像检索(DSMM-TBIR)数据引擎,并利用该引擎构建了安全多匹配TBIR(SecMM-TBIR)基准,其包含5万张监控图像与200个综合查询。此外,我们观察到,特定领域中的普通对比学习存在严重的假阴性问题,迫使模型将语义相似对推开,从而降低检索性能。为此,我们提出语义感知微调(SAFT)框架,以解决特定领域的语义压缩问题,该框架结合了语义感知软标签监督(SASS)与模态内结构蒸馏(ISD),为特定领域TBIR任务建立了有前景的范式。在多种类CLIP模型上的实验表明,与标准图像-文本对比(ITC)微调相比,SAFT在SecMM-TBIR上平均提升7.8个点的mAP@20,同时也改善了通用领域的性能。整个基准将被发布以推动进一步研究。
英文摘要:
Driven by the rapid advancement of vision-language representation learning, Text-based Image Retrieval (TBIR) has made notable progress. However, existing benchmarks are predominantly constructed on an exclusive single-match assumption between query and images. While effective in general scenarios, this assumption fails to reflect practical system performance in specific domains (e.g., surveillance), where a single query often corresponds to multiple relevant candidate images. To address this limitation, we design a Domain-Specific Multi-Match Text-based Image Retrieval (DSMM-TBIR) data engine. Leveraging this engine, we construct Security Multi-Match TBIR (SecMM-TBIR), a benchmark comprising 50k surveillance images with 200 comprehensive queries. Furthermore, we observe that vanilla contrastive learning in specific domains suffers from severe false negatives, forcing the model to push apart semantically similar pairs and thus degrading retrieval performance. We propose the Semantic-Aware Fine-Tuning (SAFT) framework to address semantic compression in specific domains, which incorporates Semantic-Aware Soft-Label Supervision (SASS) and Intra-modal Structural Distillation (ISD) to establish a promising paradigm for domain-specific TBIR tasks. Experiments across diverse CLIP-like models demonstrate that SAFT yields an average mAP@20 gain of 7.8 points on SecMM-TBIR over standard image-text contrastive (ITC) fine-tuning, while also improving general-domain performance. The entire benchmark will be released to facilitate further research.