面向细粒度图像检索的区域感知CLS令牌增强
Region-Aware CLS Token Augmentation for Fine-Grained Image Retrieval
浏览论文内容
中文总结 AI 辅助
该研究针对细粒度图像检索任务,提出区域感知CLS令牌增强方法,利用DINOv2-reg模型的寄存器令牌,自动生成局部ROI令牌并整合到ColBERT式多向量框架,提升检索性能且适配大规模搜索。
中文摘要 AI 辅助
图像检索方法通常依赖于从图像中提取的单个全局语义描述符,例如视觉Transformer中的[CLS]令牌。然而,试图将图像的所有语义信息压缩到单个描述符中会损害下游检索性能,尤其是对于细粒度检索任务。在这项工作中,我们对新型视觉Transformer中的语义令牌——全局[CLS]令牌和四个寄存器令牌(register tokens)进行增强,添加精心选择的空间令牌集合,旨在捕获表征每个语义令牌所捕获内容的空间区域表示。我们利用DINOv2-reg模型,该模型包含能自发学习对象和基于部件表示的寄存器令牌。对于每个“提示”令牌([CLS]和每个寄存器令牌),我们找到一个“伙伴”图像块令牌并提取N×N块区域以生成一组局部ROI令牌。我们的方法无需任何外部边界框或显著性模块,仅通过匹配语义令牌与其空间表示区域即可自动捕获重要的感兴趣区域。此外,我们将这些令牌整合到受ColBERT启发的多向量检索框架中,通过逐令牌对齐机制实现细粒度匹配,同时避免存储所有块嵌入带来的高存储成本。通过大量实验,我们发现:(1)寄存器令牌编码可补充[CLS]令牌的有用细粒度细节;(2)自动池化的ROI令牌进一步提升细粒度区分能力;(3)使用少量令牌的多向量检索优于DINOv2-reg单向量基线,同时对于大规模搜索仍具有可处理性。代码可在该https URL获取。
英文摘要
Image retrieval methods often rely on a single global semantic descriptor extracted from an image, e.g., the [CLS] token in vision transformers. However, trying to squeeze all the semantic information of an image into a single descriptor can hurt downstream retrieval performance, especially for fine-grained retrieval tasks. In this work, we augment the semantic tokens in the newer visual transformers, the global [CLS] token and the four register tokens, with a carefully selected collection of spatial tokens, aiming to capture the spatial region representation that characterizes the contents captured in each of the semantic tokens. We leverage the DINOv2-reg model, which includes register tokens that emergently learn object and part-based representations. For each "cue" token ([CLS] and each register token), we find a "buddy" image patch token and extract an N x N patch region to produce a set of localized ROI tokens. Our approach automatically captures important regions of interest without any external bounding boxes or saliency modules, purely by matching semantic tokens with their spatial representation regions. Furthermore, we incorporate these tokens into a multi-vector retrieval framework inspired by ColBERT, enabling fine-grained matching via a per-token alignment mechanism while avoiding the large storage cost of keeping all patch embeddings. Through extensive experiments, we find that (1) register tokens encode useful fine-grained details that can complement the [CLS] token; (2) automatically pooled ROI tokens further improve fine-grained discrimination; and (3) multi-vector retrieval with a small set of tokens improves over a DINOv2-reg single-vector baseline while remaining tractable for large-scale search. The code is available at https://github.com/IdhcbIan/Augmenting_CLS_with_ROI_tokens.
发表机构
- Institute of Mathematics and Computer Science (ICMC), University of São Paulo (USP)(圣保罗大学数学与计算机科学研究所)
- Temple University(天普大学)
机构由 AI 辅助整理,请以论文原文为准。