发表机构
CASIA; Alibaba Group; University of Chinese Academy of Sciences; Yale University(中国科学院自动化研究所; 阿里巴巴集团; 中国科学院大学; 耶鲁大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
UEmbed是一种仅解码器的多模态嵌入模型,可单次前向传播生成稀疏与稠密表示,在MMEB-v2和BEIR上表现优异,统一了两类嵌入并支持多模态智能体应用。
AI 中文摘要
稀疏检索是现代搜索系统的基础,从网页搜索到检索增强生成均是如此。现有研究引入了学习型稀疏检索(LSR),以超越精确词汇匹配,实现更丰富的语义。然而,LSR至今仍与编码器式双向架构绑定,其向多模态场景的扩展仍严重依赖辅助跨模态模块。为解决这些局限,我们提出UEmbed(统一嵌入),这是一种仅解码器的多模态嵌入模型,可在一次因果前向传播中同时生成稀疏词汇和稠密表示。UEmbed会在输入中附加N个可学习特殊标记,并将词汇表划分为N个不相交子集。每个标记的因果隐藏状态会预测其分配子集上的稀疏权重,且N个子集会被拼接为完整的稀疏向量。在公开数据上训练后,我们发布了规模为2B、4B和9B的UEmbed。UEmbed-9B在MMEB-v2上的稠密指标达到71.8、稀疏指标达到71.0,优于在公开数据上训练的多模态嵌入模型(如RzenEmbed)。在BEIR上,UEmbed也与强大的稠密和稀疏基准保持竞争力。此外,我们从有效性、效率和智能体应用三个维度展示了UEmbed的实用价值。总体而言,UEmbed提供了一种新范式:它在单个模型中统一了稠密和稀疏嵌入,同时进一步扩展了稀疏检索以统一文本和多模态输入。
英文摘要
Sparse retrieval underpins modern search systems, from web search to retrieval-augmented generation. Existing work has introduced Learned Sparse Retrieval (LSR) to push beyond exact lexical matching toward richer semantics. Yet LSR has so far remained tied to encoder-style bidirectional architectures, and its extension to multimodal settings still relies heavily on auxiliary cross-modal modules. To address these limitations, we introduce UEmbed (Unified Embedding), a decoder-only multimodal embedding model that produces both sparse lexical and dense representations in one causal forward pass. UEmbed appends N learnable special tokens to the input and partitions the vocabulary into N disjoint subsets. Each token's causal hidden state predicts sparse weights over its assigned subset, and the N subsets are concatenated into the full sparse vector. Trained on public data, we release UEmbed at 2B, 4B, and 9B scales. UEmbed-9B reaches 71.8 (dense) and 71.0 (sparse) on MMEB-v2, outperforming multimodal embedding models trained on publicly available data (e.g., RzenEmbed). On BEIR, UEmbed also remains competitive with strong dense and sparse baselines. Furthermore, we demonstrate the practical utility of UEmbed across three dimensions: effectiveness, efficiency, and agentic applications. Overall, UEmbed offers a new paradigm: it unifies dense and sparse embeddings in one model, while further extending sparse retrieval to unify text and multimodal inputs.