SetMIR:将多兴趣检索视为集合预测问题
SetMIR: Multi-Interest Retrieval as Set Prediction
浏览论文内容
中文总结 AI 辅助
SetMIR将多兴趣检索视为集合预测,用Transformer编码用户行为历史,经实验在Snap DPA数据上表现优于现有方法,部署后显著提升推荐系统的CTR与CVR。
中文摘要 AI 辅助
基于嵌入的检索是工业推荐系统的核心,但单一用户嵌入往往过于局限,无法捕捉用户的多样兴趣。多兴趣检索通过使用多个用户嵌入解决这一问题,但现有方法仍存在两个问题:兴趣崩溃(不同嵌入学习到相同兴趣)和静态分配(即使部分嵌入不必要,服务阶段仍使用固定的检索预算)。我们提出SetMIR,将多兴趣检索视为集合预测问题。SetMIR使用Transformer编码用户的行为历史,并利用K个可学习查询解码一组用户兴趣,每个兴趣生成一个检索嵌入和一个存在分数。训练期间,匈牙利匹配为查询一对一分配目标,使匹配的查询学习不同兴趣,同时存在头学习哪些查询是活跃的。服务阶段,SetMIR利用存在分数和查询级别的非极大值抑制(NMS),仅发出活跃、非冗余的近似最近邻(ANN)查询。在Snap的动态产品广告(DPA)数据上,SetMIR在所有指标上均优于四种学习型多兴趣检索器,且每个请求发出的ANN查询减少33%。作为新的检索源部署在DPA生产栈中后,SetMIR使整体转化率(CVR)提升3.1%,相较于使用相同物品嵌入、ANN索引和检索配额的物品到物品检索源,点击率(CTR)提升44%、CVR提升51%。
英文摘要
Embedding-based retrieval is at the core of industrial recommender systems, but a single user embedding is often too limited to capture a user's diverse interests. Multi-interest retrieval addresses this by using multiple user embeddings, yet existing methods still suffer from two issues: interest collapse, where different embeddings learn the same interest, and static dispatch, where serving uses a fixed retrieval budget even when some embeddings are unnecessary. We propose SetMIR, which treats multi-interest retrieval as a set prediction problem. SetMIR encodes a user's behavior history with a transformer and uses K learnable queries to decode a set of user interests, each producing a retrieval embedding and a presence score. During training, Hungarian matching assigns targets to queries one-to-one, so matched queries learn distinct interests and the presence head learns which queries are active. At serving time, SetMIR uses presence scores and query-level Non-Maximum Suppression (NMS) to issue only active, non-redundant ANN queries. On Snap's Dynamic Product Ads (DPA) data, SetMIR outperforms four learned multi-interest retrievers on every metric while issuing 33% fewer ANN queries per request. Deployed as a new retrieval source in the DPA production stack, SetMIR lifts overall CVR by 3.1%, while lifting CTR by 44% and CVR by 51% over the item-to-item retrieval source with the same item embeddings, ANN index, and retrieval quota.
发表机构
- Snap Inc.(Snap公司)
机构由 AI 辅助整理,请以论文原文为准。