arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.30251cs.IR

SetMIR:将多兴趣检索视为集合预测问题

SetMIR: Multi-Interest Retrieval as Set Prediction

Xiaodong Liu, Congfei Zhang, Hsiang-wei Chao, Siman Wang, Xiao Bai, Tong Zhao, Jingxiao Ma, Wen Zhang, Zhe Liu, Shantanu Aggarwal, Di Huang, William Leach, Yunz… 展开作者

Xiaodong Liu, Congfei Zhang, Hsiang-wei Chao, Siman Wang, Xiao Bai, Tong Zhao, Jingxiao Ma, Wen Zhang, Zhe Liu, Shantanu Aggarwal, Di Huang, William Leach, Yunzhi Zhou, Yajun Wang, Jinchao Li, Yu Zhang

首次发表
浏览论文内容

中文总结 AI 辅助

SetMIR将多兴趣检索视为集合预测,用Transformer编码用户行为历史,经实验在Snap DPA数据上表现优于现有方法,部署后显著提升推荐系统的CTR与CVR。

中文摘要 AI 辅助

基于嵌入的检索是工业推荐系统的核心,但单一用户嵌入往往过于局限,无法捕捉用户的多样兴趣。多兴趣检索通过使用多个用户嵌入解决这一问题,但现有方法仍存在两个问题:兴趣崩溃(不同嵌入学习到相同兴趣)和静态分配(即使部分嵌入不必要,服务阶段仍使用固定的检索预算)。我们提出SetMIR,将多兴趣检索视为集合预测问题。SetMIR使用Transformer编码用户的行为历史,并利用K个可学习查询解码一组用户兴趣,每个兴趣生成一个检索嵌入和一个存在分数。训练期间,匈牙利匹配为查询一对一分配目标,使匹配的查询学习不同兴趣,同时存在头学习哪些查询是活跃的。服务阶段,SetMIR利用存在分数和查询级别的非极大值抑制(NMS),仅发出活跃、非冗余的近似最近邻(ANN)查询。在Snap的动态产品广告(DPA)数据上,SetMIR在所有指标上均优于四种学习型多兴趣检索器,且每个请求发出的ANN查询减少33%。作为新的检索源部署在DPA生产栈中后,SetMIR使整体转化率(CVR)提升3.1%,相较于使用相同物品嵌入、ANN索引和检索配额的物品到物品检索源,点击率(CTR)提升44%、CVR提升51%。

英文摘要

Embedding-based retrieval is at the core of industrial recommender systems, but a single user embedding is often too limited to capture a user's diverse interests. Multi-interest retrieval addresses this by using multiple user embeddings, yet existing methods still suffer from two issues: interest collapse, where different embeddings learn the same interest, and static dispatch, where serving uses a fixed retrieval budget even when some embeddings are unnecessary. We propose SetMIR, which treats multi-interest retrieval as a set prediction problem. SetMIR encodes a user's behavior history with a transformer and uses K learnable queries to decode a set of user interests, each producing a retrieval embedding and a presence score. During training, Hungarian matching assigns targets to queries one-to-one, so matched queries learn distinct interests and the presence head learns which queries are active. At serving time, SetMIR uses presence scores and query-level Non-Maximum Suppression (NMS) to issue only active, non-redundant ANN queries. On Snap's Dynamic Product Ads (DPA) data, SetMIR outperforms four learned multi-interest retrievers on every metric while issuing 33% fewer ANN queries per request. Deployed as a new retrieval source in the DPA production stack, SetMIR lifts overall CVR by 3.1%, while lifting CTR by 44% and CVR by 51% over the item-to-item retrieval source with the same item embeddings, ANN index, and retrieval quota.

发表机构

  • Snap Inc.(Snap公司)

机构由 AI 辅助整理,请以论文原文为准。

↑