发表机构
Snap Inc.(Snap公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
GRP是一个将检索、排序和奖励建模统一到单一编码器-解码器模型的生成式推荐框架,通过mGRPO强化学习后训练和多项服务优化,在在线实验中实现观看时间与份额提升,支持渐进式部署。
AI 中文摘要
工业推荐系统依赖于多阶段级联架构,其检索、排序和服务组件难以联合替换。我们提出了GRP,一种生成式推荐框架,将检索、排序和奖励建模整合到单个编码器-解码器模型中,并评估了迈向端到端推荐的渐进路径。该模型生成多模态语义ID,并通过联合训练的排序模块对候选进行评分。冻结的排序模块随后为强化学习后训练提供奖励。我们引入了mGRPO,它在奖励优化中增加了参考锚定的边际,以保留记录目标的可能性。离线实验考察了历史编码、模型容量分配、事件选择、分词和奖励判别。服务优化将端到端检索延迟降低了69%。在线实验评估了该模型作为检索源、早期排序绕过以及替换较弱源的效果。在仅检索的比较中,相对于生产环境,观看时间增加了0.46%,份额增加了0.77%。另一项结合绕过和源替换的比较显示,观看时间增加了0.82%,份额增加了2.56%,平台级护栏保持中性。这些结果支持渐进式部署,同时指出了排序质量和推荐指标性能方面的剩余差距。
英文摘要
Industrial recommendation systems rely on multi-stage cascades whose retrieval, ranking, and serving components are difficult to replace jointly. We present GRP, a generative recommendation framework that combines retrieval, ranking, and reward modeling in a single encoder-decoder model, and evaluate a progressive path toward end-to-end recommendation. The model generates multimodal Semantic IDs and scores candidates with a jointly trained ranking module. The frozen ranking module then supplies rewards for reinforcement-learning post-training. We introduce mGRPO, which adds a reference-anchored margin to reward optimization to preserve the likelihood of logged targets. Offline experiments examine history encoding, model capacity allocation, event selection, tokenization, and reward discrimination. Serving optimizations reduce end-to-end retrieval latency by 69%. Online experiments evaluate the model as a retrieval source, with early-ranking bypass, and with replacement of weaker sources. In a retrieval-only comparison, view time increases by 0.46% and shares by 0.77% relative to production. A separate comparison combining bypass and source replacement yields increases of 0.82% in view time and 2.56% in shares, with neutral platform-level guardrails. These results support progressive deployment while identifying remaining gaps in ranking quality and performance across recommendation metrics.
Comments26 pages, 3 figures, 11 tables. Technical report