arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

MoEMB:利用高效混合专家模型扩展通用多模态嵌入

MoEMB: Scaling Universal Multimodal Embeddings with Efficient Mixture-of-Experts Models

Xuanming Cui, Shlok Kumar Mishra, Wentao Bao, Aashu Singh, Zihao Wang, Xiangjun Fan, Jun Xiao, Ser-Nam Lim, Jianpeng Cheng

arXiv 2609.08663首次发表:更新:

发表机构

University of Central Florida; Meta(中佛罗里达大学; Meta)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对通用多模态嵌入扩展难的问题,提出MoEMB,通过混合专家沿专家轴扩展编码器容量,以3B激活参数超越4倍参数的TTE方法,并首次系统研究自适应计算以提升效率。

AI 中文摘要

通用多模态嵌入(UME)日益需要编码器具备处理更广泛任务和更复杂模态的能力。先前的扩展方法要么增加表示大小、检索开销,要么将编码器扩展为重型多模态大语言模型。近期工作,如Think-Then-Embed(TTE),探索通过推理令牌进行扩展。然而,嵌入模型难以扩展:直接增加参数会与对比学习所需的大批量训练产生权衡,且检索必须在严格的延迟下进行。此外,UME任务在复杂度上具有多样性,扩展嵌入器可能带来显著的计算冗余。在本工作中,我们提出MOEMB,它通过混合专家(MoE)沿专家轴扩展UME,在保持单向量、非自回归编码的同时增加编码器容量。通过对基于MoE的UME设计空间和训练方案的系统研究,MOEMB在MMEB-V2和MRMR上,在仅使用3B激活参数的情况下,以显著更少的计算量超越了基于TTE的方法(其激活参数超过4倍),在基于公共MMEB系列数据训练的模型中树立了新的最先进水平。为进一步提升可扩展性和效率,我们首次对基于MoE的嵌入自适应计算进行了全面研究,涵盖了基于训练和仅推理的多种策略。综合这些结果,支持专家扩展作为UME有效且高效的方向,而自适应计算则进一步提升基于MLLM的嵌入模型在大规模检索和推荐系统中的效率。

英文摘要

Universal multimodal embedding (UME) increasingly demands encoder's capacity for handling a broad range of tasks and modalities with increased complexity. Prior scaling methods either increase the representation size, retrieval effort, or scales the encoder into a heavy multimodal LLM. Recent works, such as Think-Then-Embed (TTE), explore scaling via reasoning tokens. However, embedding models are hard to scale up: increasing parameters directly tradeoffs for the large training batch size that contrastive learning needs, and retrieval has to be served under tight latency. Moreover, UME tasks are diverse in complexity, where scaling up embedders can bring significant redundant computation. In this work, we propose MOEMB, which instead scales UME along the expert axis through mixture-of-experts (MoE), growing encoder capacity while preserving single-vector, non-autoregressive encoding. Through a systematic study of the design space and training recipes for MoE-based UME, MoEMB sets a new state of the art on both MMEB-V2 and MRMR among models trained on public MMEB-family data: with only 3B active parameters, MoEMB surpasses TTE-based methods with >4x active parameters, using significantly less computes. To further improve the scalability and efficiency, we conduct the first comprehensive study of adaptive computation for MoE-based embedding, spanning diverse strategies across training-based and inference-only methods. Together, these results support expert scaling as an effective and efficient direction for UME, with adaptive computation further improving efficiency for MLLM-based embedding models towards large-scale retrieval and recommendation systems.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑