全交互通用嵌入器
Omni-Interactive Universal Embedder
- Sony Group Corporation(索尼集团公司)
- Sony AI(索尼人工智能)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
该研究提出OmniUE,一种支持文本、视频、音频全交互查询的通用嵌入器,在多个基准上超越现有模型,引入OmniCHOIR基准验证其能力,推动全模态表示学习发展。
AI中文摘要:
多模态表示学习正从传统双塔架构转向基于大语言模型(LLM)的嵌入器,因其具备强大的指令遵循能力。尽管取得了这一进展,现有方法主要聚焦于语言和图像模态,这些仍是当前嵌入器中用户条件交互的主导模态。本文提出首个全交互通用嵌入器(Omni-Interactive Universal Embedder,OmniUE),它不仅通过利用专用可学习令牌的中间层表示,学习跨文本、视频和音频的统一嵌入空间,还支持全交互查询,使用户能够以文本、视觉感兴趣区域和音频片段的形式提供输入。在OmniUE内部,视觉和音频分割器处理多样化的用户交互,并将其与全模态LLM结合,通过上下文聚合生成用户条件下的任意到任意嵌入。为评估OmniUE的全交互能力,我们引入OmniCHOIR,这是一个基于给定文本、视频、音频以及单模态或多模态交互提示的全交互组合音频检索基准。OmniUE在不同模态上始终优于现有最先进的基线模型:在文本交互视频基准(MMEB-v2-video)上平均提升10.5%,在音频任务(MAEB)上提升1.1%,在视觉交互基准(SCaR)上提升83.7%,在我们的全交互OmniCHOIR基准上提升24.1%。我们认为,共同推进全模态表示学习和全交互查询,将为通用嵌入器铺平道路。
英文摘要:
Multimodal representation learning has been shifting from traditional two-tower architectures to large language model (LLM)-based embedders due to their strong instruction-following capabilities. Despite this progress, existing approaches primarily focus on language and image modalities, which also remain the dominant modalities for user-conditioned interactions in current embedders. In this paper, we propose the first Omni-Interactive Universal Embedder (OmniUE), which not only learns a unified embedding space across text, video, and audio by leveraging intermediate-layer representations from dedicated learnable tokens, but also supports omni-interactive querying, enabling users to provide inputs in the form of text, visual regions of interest, and audio spans. Within OmniUE, visual and audio segmenters process diverse user interactions and integrate them with an omni-LLM to produce user-conditioned any-to-any embeddings via context aggregation. To evaluate OmniUE's omni-interactive capabilities, we introduce OmniCHOIR, benchmarking models for omni-interactive compositional audio retrieval based on the given text, video, and audio as well as unimodal or multimodal interaction prompts. OmniUE consistently surpasses state-of-the-art baselines across diverse modalities, with average improvements of 10.5% on textual-interactive video benchmarks (MMEB-v2-video), 1.1% on audio tasks (MAEB), 83.7% on visual-interactive benchmarks (SCaR), and 24.1% on our omni-interactive OmniCHOIR benchmark. We believe that jointly advancing omni-modal representation learning and omni-interactive querying paves the way toward universal embedders.