arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.12987cs.IRcs.AI

基于双角色标识符的生成式通用多模态检索

Generative Universal Multimodal Retrieval with Dual-role Identifiers

Kaipeng Li, Haitao Yu, Xuanchen Zhou

首次发表
浏览论文内容

中文总结 AI 辅助

本文针对生成式信息检索的三大挑战,提出具备双角色标识符的DrIG框架,经实验验证其在多模态检索任务中性能优于现有基线,且能平衡效率与效果。

中文摘要 AI 辅助

生成式信息检索(GIR)作为传统“索引-检索-排序”检索流程的替代方案,通过训练生成器直接生成相关项目的标识符,已成为极具吸引力的技术路线。尽管该技术前景广阔,但仍存在诸多未解决的挑战:其一,受限的自左至右解码易受前缀级错误和局部最优问题影响;其二,多数现有GIR研究仍以单模态为主,针对文本、图像及图文混合项目的指令感知检索尚未得到充分探索;其三,基于离散标识符的GIR虽效率更高,但其检索精度仍落后于最先进的密集向量检索方法。针对这些挑战,本文提出DrIG——一种具备双角色标识符的新型通用多模态检索生成式框架,支持跨多模态、多领域的多样化检索任务。每个候选项目被分配一个单一的残差量化标识符,该标识符承担两个互补角色:在序列角色中,标识符以自回归方式解码,第一个token明确建模模态,其余token捕获逐步细化的语义;在集合角色中,相同token被重新解释为无序集合,以提供与前缀无关的相关性先验,指导受限束搜索并缓解局部最优错误。在M-BEIR基准及文本到图像评估数据集上开展的大量实验表明:(1)DrIG在各类任务中始终优于最先进的生成式多模态基线,而混合重排序在与强大密集检索器的对比中实现了效率与效果的良好平衡;(2)消融实验与缩放分析揭示了基础大型多模态模型(LMM)、束大小、重排序深度及融合策略对检索性能的影响,为系统设计提供了实用指导。

英文摘要

Generative information retrieval (GIR) has emerged as a compelling alternative to the conventional index-retrieve-then-rank retrieval pipeline by training a generator to produce the identifiers of relevant items directly. Despite its promise, a number of open challenges still remain. First, constrained left-to-right decoding is vulnerable to prefix-level errors and local optima. Second, most prior GIR research remains largely unimodal, leaving instruction-aware retrieval across text, image, and mixed image-text items underexplored. Third, although discrete identifier-based GIR offers higher efficiency, its retrieval accuracy still lags behind that of the cutting-edge dense-vector-based retrieval methods. Motivated by these challenges, we propose DrIG, a novel Generative framework for universal multimodal retrieval featuring Dual-role Identifiers, which supports diverse retrieval tasks across multiple modalities and domains. Each candidate is assigned a single residual-quantized identifier that serves two complementary roles. In its sequential role, the identifier is decoded autoregressively, where the first token explicitly models modality and the remaining tokens capture progressively finer semantics. In its set-based role, the same tokens are reinterpreted as an unordered set to provide a prefix-independent relevance prior, which guides constrained beam search and alleviates local-optimum errors. Extensive experiments on the M-BEIR benchmark and the text-to-image evaluation datasets show that:(1)DrIG consistently outperforms state-of-the-art generative multimodal baselines across diverse tasks, while hybrid reranking achieves a favorable efficiency-effectiveness trade-off against strong dense retrievers. (2)Ablation and scaling analyses reveal how the base LMM, beam size, reranking depth, and fusion strategy affect retrieval performance, providing practical guidance for system design.

发表机构

  • Institute of Library, Information and Media Science, University of Tsukuba(筑波大学图书馆、信息与媒体科学研究所)
  • College of Knowledge and Library Sciences, University of Tsukuba(筑波大学知识与图书馆科学学院)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑