arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

概念补充稠密语义:学习用于文本-图像检索的紧凑稀疏空间

Concepts Complement Dense Semantics: Learning Compact Sparse Spaces for Text-Image Retrieval

Yoonseo Kim, Jungwoo Choi, Cheonyoung Park, Youngwook Kim, Yongho Song, SeongKu Kang

arXiv 2609.32671首次发表:更新:

发表机构

Korea University; Sungkyunkwan University; KT Corporation(高丽大学; 成均馆大学; KT公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出GRASP框架,通过挖掘视觉-文本概念并训练轻量级稀疏头,生成可解释的概念级证据以补充稠密语义匹配,在文本-图像检索中实现更优准确率与更紧凑接地的稀疏空间。

AI 中文摘要

跨模态检索已通过将图像和文本编码到共享稠密嵌入空间的视觉-语言预训练模型得到推进。虽然稠密表示能有效捕获整体语义相似性,但它们常常掩盖了精确跨模态匹配所需的细粒度视觉-文本信息。近期方法引入学习得到的稀疏分支,以词汇证据补充稠密匹配,但它们依赖冗余的语言模型词元空间,且缺乏对稀疏维度的显式接地。我们提出GRASP,一个紧凑且接地的稀疏学习框架,从语料库中挖掘视觉-文本概念。一个轻量级稀疏头被训练来预测与每个图像或文本相关的概念,产生可解释的概念级证据,以补充稠密语义匹配。大量实验表明,GRASP在检索准确性上优于最先进的稠密-稀疏基线,同时产生更紧凑且接地的稀疏空间。

英文摘要

Cross-modal retrieval has been advanced by vision-language pre-trained models that encode images and texts into a shared dense embedding space. While dense representations effectively capture overall semantic similarity, they often obscure fine-grained visual-textual information needed for precise cross-modal matching. Recent methods introduce a learned sparse branch to complement dense matching with lexical evidence, but they rely on a redundant language-model token space and lack explicit grounding for sparse dimensions. We propose GRASP, a compact and grounded sparse learning framework that mines visual-textual concepts from the corpus. A lightweight sparse head is trained to predict concepts relevant to each image or text, yielding interpretable concept-level evidence that complements dense semantic matching. Extensive experiments show that GRASP improves retrieval accuracy over the state-of-the-art dense-sparse baselines while yielding a more compact and grounded sparse space.

CommentsAccepted for oral presentation at KEIR@CIKM 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑