arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Omni-Embed-Mini:通过密集蒸馏在不遗忘的情况下绑定模态

Omni-Embed-Mini: Binding Modalities Without Forgetting via Dense Distillation

Mohammed Irfan Kurpath, Jaseel Muhammad Kaithakkodan, Sahal Shaji Mullappilly, Ivan Laptev, Hisham Cholakkal

arXiv 2610.02148首次发表:更新:

发表机构

Mohamed Bin Zayed University of Artificial Intelligence (MBZUAI)(穆罕默德·本·扎耶德人工智能大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出Omni-Embed-Mini,通过密集蒸馏和共享骨干网络,在不更新文本参数的情况下将多模态嵌入对齐到统一空间,保持文本检索性能并显著减小模型规模。

AI 中文摘要

将文本嵌入模型扩展到新模态通常会导致文本检索质量下降,现有的全模态嵌入器通过数十亿参数来弥补这一问题。我们提出了Omni-Embed-Mini,一个0.9B参数的模型,它将文本、语音、音频、图像、视频和富含视觉的文档映射到一个共享的余弦空间中,且不更新任何文本侧参数。我们的关键见解是,教师信号不需要单独的嵌入模型:每个媒体样本都配有一个密集的级联标题,教师目标仅仅是冻结骨干网络对该标题的嵌入。由于教师和学生共享相同的骨干权重,它们处于字节相同的几何空间中,轻量级投影仪和模态编码器上的分阶段LoRA适配器足以实现对齐。训练结合了Matryoshka SigLIP对比损失和一个在线混合难负样本挖掘器,其负样本随着编码器的改进而变得更加困难。该方案通过替换为原生视觉-语言骨干网络,可扩展到2.3B变体。Omni-Embed-Mini-0.9B保持其文本权重与骨干网络逐位相同,因此训练不会使文本检索退化(在MTEB-v2 BEIR-8上nDCG@10为49.57),同时将其扩展到五种额外的模态,并且比我们比较的所有开放全模态嵌入器小约2.7倍至9.5倍。2.3B变体与封闭的gemini-embedding-2具有竞争力,在整体模态平均上略微领先。模型、代码、数据和评估工具在我们的项目页面上:此https URL

英文摘要

Extending a text embedding model to new modalities typically degrades text retrieval quality, and existing omni-modal embedders compensate with multi-billion parameters. We present Omni-Embed-Mini, a 0.9B-parameter model that maps text, speech, audio, images, video, and visually-rich documents into a single shared cosine space without updating any text-side parameter. Our key insight is that the teacher signal requires no separate embedding model: each media sample is paired with a dense cascaded caption, and the teacher target is simply the frozen backbone's own embedding of that caption. Because teacher and student share the same backbone weights, they inhabit byte-identical geometry, and lightweight projectors plus phased LoRA adapters on the modality encoders suffice for alignment. Training combines a Matryoshka SigLIP contrastive loss with an online hybrid hard-negative miner whose negatives sharpen as the encoder improves. The recipe carries over to a 2.3B variant by swapping in a native vision-language backbone. Omni-Embed-Mini-0.9B keeps its text weights bit-identical to the backbone, so training cannot regress text retrieval (49.57 nDCG@10 on MTEB-v2 BEIR-8), while extending it to five additional modalities, and is ~2.7x to 9.5x smaller than every open omni embedder we compare against. The 2.3B variant is competitive with the closed gemini-embedding-2, edging ahead of it on the overall-modality average. Models, code, data and evaluation harness are on our project page: https://omniembed.cvmbzuai.com

CommentsFindings of EMNLP 2026. 26 pages, 8 figures, 14 tables. Project page: https://omniembed.cvmbzuai.com

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑