发表机构
University of Technology Sydney; New York University Abu Dhabi; University of California, San Diego(悉尼科技大学; 纽约大学阿布扎比分校; 加利福尼亚大学圣地亚哥分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出SPRINT,一种基于平均概率速度的单步生成式推荐方法,通过双向Transformer在单次前向传播中生成SID,并设计双层对比目标保持令牌连贯性,实现高效且准确的推荐。
AI 中文摘要
基于语义ID(SID)的生成式推荐将每个条目表示为离散令牌序列,并通过生成用户可能交互的条目的SID来进行推荐。该领域中的两种主流范式通常逐令牌生成:自回归模型从左到右解码令牌,而非自回归模型并行解码,但仍需多轮精炼以保持竞争力。因此,两者通常每个条目需要多次前向传播,这一成本在延迟敏感型推荐系统中是难以承受的。我们提出一个问题:能否在单次前向传播中生成一个条目?并通过一个我们称之为“平均概率速度”的新视角来回答这一问题。我们将SID生成视为令牌生成概率的流动,并用整个生成过程中的平均速度来刻画它。我们证明,这一平均速度完全由每个令牌的平均生成概率决定。因此,我们直接使用双向Transformer在单次前向传播中参数化并学习所有令牌的概率。由于这些概率是跨位置独立生成的,令牌之间的连贯性会丢失,我们进一步设计了一种双层流动对比目标来恢复条目令牌间的连贯性。该目标在令牌级别和SID级别上将目标SID与负SID进行对比。令牌级别将目标令牌的生成概率排在负SID的令牌之上,而SID级别则将每个SID的令牌作为一个整体条目进行评分,以捕捉每个条目的令牌连贯性。大量实验表明,我们的模型不仅生成推荐的效率远高于其他方法(比第二快的自回归/非自回归方法快8.39-10.04倍),而且推荐准确性也更优(比次优方法平均提升7.77%)。
英文摘要
Semantic ID (SID) based generative recommendation represents each item as a sequence of discrete tokens, and recommends by generating the SID of the item a user would like to interact with. Both dominant paradigms in this domain generally pay for generation token by token: autoregressive models decode the tokens left-to-right, while non-autoregressive models decode in parallel yet still need multiple rounds of refinement to stay competitive. Therefore, both generally spend multiple forward passes per item, a cost that is prohibitive in latency-sensitive recommender systems. We ask whether an item can be generated in a single forward pass, and answer it through a new perspective which we call average probability velocity. We view SID generation as a flow of token generation probabilities and characterize it by its average velocity over the whole generation process. We prove that this average velocity is fully determined by the average generation probability of each token. Therefore, we directly parameterize and learn the probabilities of all tokens in a single forward pass with a bidirectional Transformer. As these probabilities are generated independently across positions and the coherence among tokens is lost, we further design a dual-level flow contrastive objective to restore the coherence among an item's tokens. It contrasts the target SID against negative SIDs at both the token and SID levels. The token level ranks the generation probabilities of the target tokens above those of negative SIDs, while the SID level scores the tokens of each SID as a whole item for capturing token coherence of each item. Extensive experiments show that our model not only generates recommendations far more efficiently ($8.39-10.04\times$ speedup over the second-fastest AR/NAR method) but also attains superior recommendation accuracy ($7.77\%$ average improvement over the second-best.