arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

抖音多模态嵌入模型技术报告

Douyin Multimodal Embedding Model Technical Report

Haonan Chen, Chu Li, Zhicheng Wang, Yuanwei Liu, Yuanjiang Wang, Shaohua Jiang, Zhicheng Dou

arXiv 2608.02148首次发表:更新:

发表机构

ByteDance; Renmin University of China(字节跳动; 中国人民大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对现有多模态嵌入模型难以兼顾效率与细粒度区分的问题,提出分两阶段训练的DME模型,在MMEB-v2数据集及抖音生产场景中均取得优异效果。

AI 中文摘要

多模态表示学习是现代人工智能的基石,通过将多模态查询与目标编码为向量,为工业级搜索、推荐提供动力,并为现代智能体提供支撑。拥有复杂模态与大规模内容的现实平台,如抖音、小红书、YouTube,既需在十亿级索引规模下具备高效性,又需在困难匹配中具备细粒度区分能力。现有多模态大语言模型(MLLM)嵌入模型极少能同时满足这两点:对比模型效率高,但依赖的对级监督对细粒度区分而言过于粗糙;基于思维链(CoT)的模型虽通过显式生成提升了区分能力,但不适合在线服务。本文提出抖音多模态嵌入模型(DME),该模型分两个阶段训练以结合两者优势:第一阶段执行大规模对比预训练,构建覆盖广泛模态与任务的统一多模态嵌入空间;第二阶段通过两种机制补充语义充分性,即嵌入基于检索相关证据并保留细粒度对侧语义的特性:证据 grounded 类型隐式推理通过隐空间内的隐式推理组织检索证据,跨条件重建通过跨方向自回归重建强化对侧语义。两种机制仅在训练阶段生效,仅增加极少查询侧开销,因此 DME 的服务效率与标准对比编码器相当。在 MMEB-v2 数据集上,DME 的 2B 和 9B 变体在相当规模下达到了最优结果,分别为 74.8 和 78.4,在视频及视觉-文档任务上表现尤为突出。在生产环境中,DME 在抖音内部离线评估集上实现了 2.92% 的相对提升,已部署在抖音生成式、图像、AI 搜索等场景中,并在抖音搜索的在线 A/B 测试中获得了 0.1% 的终身(LT)提升。

英文摘要

Multimodal representation learning is a cornerstone of modern AI. By encoding multimodal queries and targets into vectors, it powers industrial search and recommendation and underpins modern agents. Real-world platforms with complex modalities and massive-scale content, such as Douyin, Xiaohongshu, and YouTube, demand both efficiency under billion-scale indexing and fine-grained discrimination for hard matching. Existing MLLM embedding models rarely satisfy both. Contrastive models are efficient but rely on pair-level supervision too coarse for fine-grained distinctions, while CoT-based models improve discrimination through explicit generation impractical to serve online. We present Douyin Multimodal Embedding (DME), a model trained in two stages to combine both strengths. Stage 1 performs large-scale contrastive pre-training that establishes a unified multimodal embedding space with broad modality and task coverage. Stage 2 supplements semantic sufficiency, the property that an embedding is grounded in retrieval-relevant evidence and preserves fine-grained counterpart-side semantics, via two mechanisms. Evidence-Grounded Typed Latent Reasoning organizes retrieval evidence through hidden-space latent reasoning, and Cross-Conditional Reconstruction enforces counterpart-side semantics through cross-directional autoregressive reconstruction. Both act only during training and add only marginal query-side overhead, so DME serves as efficiently as a standard contrastive encoder. On MMEB-v2, DME reaches state-of-the-art results at comparable scales for its 2B and 9B variants (74.8 and 78.4), with especially strong video and visual-document tasks. In production, DME delivers a 2.92% relative gain on Douyin's in-house offline evaluation set, is deployed across Douyin scenarios such as generative, image, and AI search, and yields a 0.1% Lifetime (LT) gain in online A/B testing on Douyin search.

CommentsTechnical Report

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑