发表机构
Shanghai Jiao Tong University; Alibaba Group(上海交通大学; 阿里巴巴集团)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对通用多模态嵌入缺乏视觉身份判别能力的问题,提出视觉身份判别统一形式化,构建MVEB基准,设计身份感知采样的学习框架,提升了UMEs的身份判别能力。
AI 中文摘要
通用多模态嵌入(Universal Multimodal Embeddings, UMEs)旨在将多种模态和任务统一到共享表示空间中。近年来,多模态大语言模型(Multimodal Large Language Models, MLLMs)的发展推动该领域取得了显著进展。然而,现有UMEs方法中,视觉身份判别这一对实例检索、重识别、AI生成内容的身份保留等任务至关重要的能力仍未得到充分探索。为填补这一空白,我们提出了视觉身份判别(VisID)的统一形式化,并引入MVEB(Multimodal Visual Identity Embedding Benchmark)——一个由真实世界和合成数据集整理而成的大规模基准,以支持评估与训练。此外,我们提出了一种简单却有效的学习框架,通过精心设计的身份感知采样机制,联合优化通用多模态和视觉身份表示。大量实验表明,我们的方法成功赋予UMEs强大的身份判别能力,同时保持了具有竞争力的通用多模态性能。我们认为,本研究不仅揭示了这一关键却被忽视的能力,还为构建更全面的通用多模态嵌入迈出了一步。代码和数据可在MVEB获取。
英文摘要
Universal Multimodal Embeddings (UMEs) aim to unify various modalities and tasks into a shared representation space. In recent years, this field has witnessed substantial progress driven by the development of Multimodal Large Language Models (MLLMs). However, a crucial capability, visual identity discrimination, remains underexplored in existing UME methods, despite its critical role in a wide range of tasks, including instance retrieval, re-identification, and identity preservation in AI-generated content. To bridge this gap, we propose a unified formulation for visual identity discrimination~(VisID) and introduce $\textbf{MVEB}$ ($\textbf{M}$ultimodal $\textbf{V}$isual Identity $\textbf{E}$mbedding $\textbf{B}$enchmark), a large-scale benchmark curated from both real-world and synthetic datasets to support evaluation and training. Furthermore, we present a simple yet effective learning framework that jointly optimizes general multimodal and visual identity representations through a carefully designed identity-aware sampling mechanism. Extensive experiments demonstrate that our approach successfully endows UMEs with strong identity discrimination capability and maintains competitive general multimodal performance. We believe this work not only illuminates a critical yet neglected capability, but also takes a step toward more holistic universal multimodal embeddings. Code and data are available at \href{https://chrisclear3.github.io/MVEB}{MVEB}.
CommentsAccepted to CVPR 2026