发表机构
Snap Inc.(Snap公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对Snap动态产品广告检索的挑战,提出CAMIE框架,基于LLM/MLLM主干与协同交互数据微调,线上部署后显著提升多场景下的CTR与CVR。
AI 中文摘要
物品到物品(I2I)检索是大规模推荐和广告系统中的核心基础任务。在生产环境的Snap动态产品广告(DPA)中,I2I检索面临两大挑战:独立的视觉、文本及多模态编码器会割裂检索栈,仅基于内容的训练无法使嵌入与驱动下游转化的协同交互行为对齐。我们提出CAMIE,一种用于Snap DPA检索的协同交互感知多模态物品嵌入框架。CAMIE基于LLM/MLLM主干构建,利用其原生多模态接口将物品图像和元数据表示在共享嵌入空间中,随后通过从用户旅程中挖掘的协同交互物品对,以对称的批量内InfoNCE目标对主干进行微调。离线实验显示,CAMIE在Recall@10指标上优于最强的商业多模态嵌入模型,且可从同一检查点提供纯文本检索,质量损失极小。在线上,CAMIE可作为两个已部署的基于内容的I2I编码器的直接替代方案,相较于多模态对照组,其点击率(CTR)提升0.390%、转化率(CVR)提升10.832%;相较于文本对照组,CTR提升18.958%、CVR提升13.12%;在整体DPA流量上,CTR提升0.211%、CVR提升1.911%,目前已部署于生产环境。
英文摘要
Item-to-item (I2I) retrieval is a core primitive in large-scale recommendation and advertising systems. In production Snap Dynamic Product Ads (DPA), I2I retrieval faces two challenges: separate visual, textual, and multimodal encoders fragment the retrieval stack, and content-only training does not align embeddings with the co-engagement behavior that drives downstream conversions. We present CAMIE, a co-engagement-aware multimodal item embedding framework for Snap DPA retrieval. CAMIE builds on LLM/MLLM backbones, using their native multimodal interfaces to represent item images and metadata in a shared embedding space. It then fine-tunes the backbone on co-engaged item pairs mined from user journeys with a symmetric in-batch InfoNCE objective. Offline, CAMIE outperforms the strongest commercial multimodal embedding model on Recall@10 and serves text-only retrieval from the same checkpoint with minimal quality loss. Online, CAMIE serves as a drop-in replacement for two deployed content-based I2I encoders, delivering +0.390% CTR / +10.832% CVR over the multimodal control, +18.958% CTR / +13.12% CVR over the text control, and +0.211% CTR / +1.911% CVR on overall DPA traffic. CAMIE is deployed in production.