Ovis-Embedding:推动通用全模态嵌入的前沿
Ovis-Embedding: Pushing the Frontiers of Universal Omni-Modal Embeddings
- Alibaba Token Hub, Alibaba Group(阿里巴巴集团阿里巴巴Token Hub)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
Ovis-Embedding通过共享多模态骨干网络原生集成文本、图像、视频和音频,采用低秩初始化、同源采样和焦点损失等优化,在多个基准上达到最先进性能,推动任意到任意检索的通用嵌入模型发展。
AI中文摘要:
在本报告中,我们介绍了Ovis-Embedding,这是一个基于文本、图像、视频和音频原生集成的最先进的全模态嵌入系列。Ovis-Embedding不是组装单独模态塔,而是使用共享的多模态骨干网络在统一的表示空间中对不同模态进行编码。具体来说,我们取得了三项关键进展:(1)原生全模态初始化:我们采用预训练的Qwen-omni模型作为嵌入骨干,并通过低秩初始化的对比训练进行适配;(2)以数据为中心的全模态训练:我们构建了一个涵盖文本、图像、视频、音频和交错多模态数据的广泛高质量语料库。为了提高数据效率,我们引入了同源采样以形成任务一致的批次,并包含信息丰富的批内负样本;(3)嵌入特定的训练和推理优化:我们使用焦点损失来强调难例,并利用基于相似性的嵌入蒸馏从互补专家中迁移细粒度相似性结构。在推理时,低秩特征分解使得嵌入紧凑且维度灵活,性能损失极小。实证评估表明,Ovis-Embedding系列在MMEB-v3、MMEB-v2、MVEB、MAEB和RTEB上取得了最先进的性能,展示了其在文本、图像、视频和音频模态上的有效性。这些结果凸显了统一全模态训练在克服模态碎片化、推进任意到任意检索的通用嵌入模型方面的潜力。
英文摘要:
In this report, we introduce \textbf{Ovis-Embedding}, a state-of-the-art omni-modal embedding family built on native integration of text, image, video, and audio. Instead of assembling separate modality towers, Ovis-Embedding uses a shared multimodal backbone to encode different modalities in a common representation space. Specifically, we make \textbf{three key advances}: (1) \textbf{native omni-modal initialization}: we adopt a pretrained Qwen-omni model as the embedding backbone and adapt it through contrastive training with low-rank initialization; (2) \textbf{data-centric omni-modal training}: we construct a broad, high-quality corpus spanning text, images, video, audio, and interleaved multimodal data. To improve data efficiency, we introduce homogeneous-source sampling to form task-consistent batches with informative in-batch negatives; and (3) \textbf{embedding-specific training and inference optimization}: we use focal loss to emphasize hard examples and similarity-based Embedding Distillation to transfer fine-grained similarity structure from complementary experts. At inference time, low-rank feature decomposition enables compact embeddings with flexible dimensionality and minimal performance loss. Empirical evaluations show that the \textbf{Ovis-Embedding} family achieves state-of-the-art performance on \textbf{MMEB-v3}, \textbf{MMEB-v2}, \textbf{MVEB}, \textbf{MAEB}, and \textbf{RTEB}, demonstrating its effectiveness across text, image, video, and audio modalities. These results highlight the potential of unified omni-modal training to overcome modality fragmentation and advance universal embedding models for any-to-any retrieval.