发表机构
Tsinghua University; Microsoft; Shanghai Jiao Tong University(清华大学; 微软; 上海交通大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对多模态检索中推理成本高的问题,提出Skim and Skip分层自适应推理框架,通过标记级证据选择和深度自适应推理,在保留99%检索性能的同时实现显著加速与FLOPs减少。
AI 中文摘要
通用多模态检索(UMR)日益采用多模态大语言模型(MLLM)作为统一嵌入主干,但它们强大的检索性能伴随着巨大的推理成本。现有方法通常依赖均匀密集推理,其中所有输入标记都通过整个模型处理,并使用最终层的[EOS]表示进行匹配。然而,这种范式忽略了多模态检索中的两种关键异质性:标记对最终检索嵌入的贡献高度不均,且不同查询需要明显不同的推理深度。为解决这一问题,我们提出Skim and Skip(SAS),一种用于高效多模态检索的分层自适应推理框架。SAS首先执行标记级证据选择,仅保留与最终检索嵌入最相关的输入信息,然后执行深度自适应推理,以确定当前表示是否已足够用于可靠匹配。在12项MMEB检索任务上的实验表明,SAS保留了约99%的密集基线平均检索性能,同时实现了高达1.64倍的端到端加速和高达66.3%的FLOPs减少。
英文摘要
Universal multimodal retrieval (UMR) increasingly adopts multimodal large language models (MLLMs) as unified embedding backbones, but their strong retrieval performance comes at substantial inference cost. Existing methods typically rely on uniformly dense inference, where all input tokens are processed through the entire model and matched using the final-layer [EOS] representation. However, this paradigm overlooks two key forms of heterogeneity in multimodal retrieval: token contributions to the final retrieval embedding are highly uneven, and different queries require markedly different amounts of inference depth. To address this, we propose Skim and Skip (SAS), a hierarchical adaptive inference framework for efficient multimodal retrieval. SAS first performs token-level evidence selection to preserve only the input information most relevant to the final retrieval embedding, and then performs depth-adaptive inference to determine whether the current representation is already sufficient for reliable matching. Experiments on 12 MMEB retrieval tasks show that SAS retains about 99% of the dense baseline's average retrieval performance while achieving up to 1.64 times end-to-end speedup and up to 66.3% FLOPs reduction.