arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

ReLoop-UME:带可学习检索寄存器的循环深度通用多模态嵌入

ReLoop-UME: Recurrent Depth with Learnable Retrieval Registers for Universal Multimodal Embedding

Shijie Wang, Xiangzhao Hao, Yueti Li, Guangyu Cao, Xinyu Tang, Haiyun Guo

arXiv 2607.28751首次发表:更新:

AI 中文总结

该研究针对通用多模态嵌入的延迟问题,提出带可学习检索寄存器的循环深度模型ReLoop-UME,在MMEB-V2等数据集上提升检索性能且运行速度显著优于现有模型。

AI 中文摘要

通用多模态嵌入(UME)将异构多模态输入映射到共享嵌入空间。现有UME模型要么通过单次前向编码形成嵌入,要么通过显式理由令牌和潜在自回归状态增加计算。尽管令牌扩展可改善复杂匹配,但串行生成会增加检索延迟,且最终嵌入依赖生成的中间状态。这引发了一个不同的问题:能否沿模型深度扩展有用计算,同时保持令牌工作空间固定?我们分析了独立训练的UME模型各层的正负相似度分离,观察到共享进展:早期层对多模态输入进行上下文化处理,连续的中晚期阶段形成检索判别特征,最终层将其映射到嵌入空间。基于这一发现,我们提出ReLoop-UME,其执行一次早期层,循环复用参数共享的检索形成模块,并在最后一次循环后应用最终映射层。可学习检索寄存器提供持久的检索特定状态,这些状态在各次循环间积累并交换证据,最终寄存器用作嵌入读出。在MMEB-V2和MRMR数据集上,ReLoop-UME在不同主干网络上的检索性能持续提升,同时比UME-R1快44.9倍,比PLUME快1.5倍。

英文摘要

Universal multimodal embedding (UME) maps heterogeneous multimodal inputs into a shared embedding space. Existing UME models either form embeddings through single forward encoding or add computation through explicit rationale tokens and latent autoregressive states. Although token expansion can improve complex matching, serial generation increases retrieval latency and makes the final embedding depend on generated intermediate states. This raises a different question: can useful computation be expanded along model depth while keeping the token workspace fixed? We analyze positive-negative similarity separation at every layer of independently trained UME models and observe a shared progression: early layers contextualize multimodal inputs, a contiguous middle-to-late stage forms retrieval-discriminative features, and the final layers map them into the embedding space. Based on this finding, we propose ReLoop-UME, which executes the early layers once, recurrently reuses a parameter-shared retrieval-forming block, and applies the final mapping layers after the last loop. Learnable Retrieval Registers provide persistent retrieval-specific states that accumulate and exchange evidence across loops, with the final register serving as the embedding readout. On MMEB-V2 and MRMR, ReLoop-UME consistently improves retrieval across different backbones while running 44.9x faster than UME-R1 and 1.5x faster than PLUME.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑