arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.04605cs.IRcs.AIcs.CLcs.CV

所有视觉令牌都同等重要吗?用于视觉语言检索的对象证据保留令牌合并

Do All Visual Tokens Matter Equally? Object-Evidence Preserving Token Merging for Vision-Language Retrieval

Suhyeong Park, Junha Jung, Jungwoo Park, Jaewoo Kang

首次发表
浏览论文内容

中文总结 AI 辅助

研究视觉语言检索中多向量方法存储和评分成本高的问题,提出对象感知令牌合并框架SaMer,压缩图像侧令牌同时保留交互接口,实验证明其能降成本且提升性能。

中文摘要 AI 辅助

多向量视觉语言检索通过最大相似性后期交互保留细粒度视觉证据,但密集图像侧令牌使存储和评分成本高昂。现有令牌压缩方法会移除或破坏对象和区域级证据。我们提出SaMer,一种对象感知令牌合并框架,在保留原始后期交互接口的同时将图像侧投影后令牌压缩为K个代表性质心。SaMer仅在训练期间使用对象注释作为合并先验,以防止跨实例混合,推理时不需要真实边界框或检测器,并且仅调整共享投影层,同时冻结视觉和语言主干。当K = 64时,SaMer移除了超过93%的图像侧令牌,并将ColPali存储减少了16.09倍,同时提高了Flickr30K和MSCOCO上的R@1。这些收益的出现是因为对象感知合并保留了查询可选择的对象证据,而修剪或仅特征池化可能会移除或破坏这些证据。SaMer还优于压缩基线,并显示出更强的短语级基础,这表明高效的多向量检索不仅取决于减少令牌数量,还取决于保留未来查询令牌需要选择的证据。

英文摘要

Multi-vector vision-language retrieval preserves fine-grained visual evidence through maximum-similarity late interaction, but dense image-side tokens make storage and scoring expensive. Existing token compression methods reduce this cost, yet they can remove or collapse object- and region-level evidence that future query tokens may need to select. We propose SaMer, an object-aware token merging framework that compresses image-side post-projector tokens into $K$ representative centroids while preserving the original late-interaction interface. SaMer uses object annotations only during training as a merge prior to discourage cross-instance mixing, requires no ground-truth bounding boxes or detectors at inference time, and adapts only the shared projection layer with frozen vision and language backbones. With $K=64$, SaMer removes more than 93% of image-side tokens and reduces ColPali storage by $16.09\times$, while improving R@1 on Flickr30K and MSCOCO. These gains arise because object-aware merging preserves query-selectable object evidence that pruning or feature-only pooling can remove or collapse. SaMer also outperforms compression baselines and shows stronger phrase-level grounding, suggesting that efficient multi-vector retrieval depends not only on reducing token count, but on preserving the evidence future query tokens need to select.

发表机构

  • The Catholic University of Korea(韩国天主教大学)
  • Korea University(韩国大学)
  • AIGEN Sciences Inc.(AIGEN科技公司)

机构由 AI 辅助整理,请以论文原文为准。

↑