arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Netflix 中通过多模态嵌入实现多媒体资产个性化

Multimedia Asset Personalization via Multimodal Embeddings at Netflix

Emma Yanyang Kong, Aditya Deshpande, Bowei Yan, Asad Abbasi, Santiago Castro, Avneesh Saluja, David Fagnan, Ashish Rastogi

arXiv 2608.18322首次发表:更新:

AI 中文总结

Netflix 采用多模态嵌入技术,通过双塔模型结合 CLIP 嵌入、三模态基础模型 MediaFM 及离线代理任务,优化了美术图像和视频预览的个性化推荐,提升了冷启动性能与效果。

AI 中文摘要

个性化推广资产(即美术图像和视频预览片段)对 Netflix 上的内容发现至关重要。传统的资产选择模型依赖基于 ID 的交互历史,无法感知资产内容,也无法为新上线的标题和资产提供服务。我们描述了多模态嵌入如何重塑 Netflix 的生产系统,并为将基础模型嵌入应用于推荐系统的从业者提供可借鉴的经验。首先,预训练图像嵌入解锁跨标题、跨画布的知识迁移。为双塔模型补充 CLIP 图像嵌入,使单个模型可服务于 Netflix 的全部五种美术画布类型,取代了五个分别训练的单画布模型,并大幅提升了冷启动性能。一个轻量扩展利用 CLIP 的联合文本-图像空间,使美术个性化在搜索中具备查询感知能力。其次,多模态在视频预览个性化方面显著优于任何单一模态。我们介绍了 MediaFM,这是我们内部训练的三模态基础模型,基于 Netflix 节目库中大规模镜头语料库训练,融合了视觉(SeqCLIP)、音频(wav2vec 2.0)和时序文本信号;将其用于视频预览个性化时,在离线测试和在线 A/B 测试中均优于强大的仅视觉基线。第三,一项性能与在线结果相关的简单离线代理任务可加速实验和产品化周期。仅从嵌入预测基于流行度的赢家,可对嵌入模型和版本进行排序,在任何端到端集成或 A/B 测试前精简选择空间;目前该任务已对每个新的 MediaFM 检查点进行筛选。我们还分享了使这些部署可行的生产工程决策(共享嵌入基础设施、低延迟服务、低成本筛选),以及遇到的设计权衡和失败模式。

英文摘要

Personalized promotional assets, namely artwork images and video preview clips, are critical to content discovery on Netflix. Traditional models for asset selection rely on ID-based interaction history, leaving them blind to asset content and unable to serve newly launched titles and assets. We describe how multimodal embeddings reshaped production systems at Netflix and report transferable lessons for practitioners adopting foundation-model embeddings into recommender systems. First, pretrained image embeddings unlock cross-title, cross-canvas knowledge transfer. Augmenting a two-tower model with CLIP image embeddings lets a single model serve all five Netflix artwork canvas types, replacing five separately trained per-canvas models and substantially improving cold-start performance. A lightweight extension reuses CLIP's joint text-image space to make artwork personalization query-aware in search. Second, multimodality decisively beats any single modality for video preview personalization. We describe MediaFM, our in-house tri-modal foundation model trained on a large-scale corpus of shots from the Netflix show catalog, fusing visual (SeqCLIP), audio (wav2vec 2.0), and timed-text signals; adopted for video preview personalization, it outperforms strong visual-only baselines both offline and in online A/B tests. Third, a simple offline proxy task whose performance correlates with online outcomes can accelerate the experimentation and productization cycle. Predicting the popularity-based winner from embeddings alone ranks embedding models and versions, pruning the choice space before any end-to-end integration or A/B test; it now gates every new MediaFM checkpoint. We also share the production engineering decisions (shared embedding infrastructure, low-latency serving, cheap screening) that made these deployments viable, along with the design tradeoffs and failure modes we encountered.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑