arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Gryphon-v2:替代级联的单一模型——带Rollout蒸馏的生成排序推荐系统

Gryphon-v2: One Model in Place of a Cascade - Generate-and-Rank Recommender with Rollout Distillation

Anna Lipkina, Daria Tikhonovich, Viktor Yanush, Mariia Ulianova, Oleg Sorokin, Vladislav Dodonov, Ilya Murzin, Denis Burshtein, Nikolay Savushkin

arXiv 2608.06213首次发表:更新:

AI 中文总结

Gryphon-v2是替代多阶段级联推荐系统的单一模型,采用生成排序架构与Rollout蒸馏技术,在Yandex Music的在线实验中使活跃用户增1.41%且服务延迟相当。

AI 中文摘要

工业推荐系统通常部署为多阶段级联结构,包含独立的候选生成器、预排序器和最终排序器。尽管这类级联结构有效,但需要重复处理用户历史、复杂的特征流水线以及多个服务阶段。基于语义ID的生成式检索为构建更简单的端到端系统提供了途径,不过仅下一个物品预测无法捕捉生产排序目标所编码的细粒度偏好。本文提出Gryphon-v2,一种用于端到端推荐的统一生成排序架构。该模型对用户历史仅编码一次,通过自回归解码器生成语义ID候选,将其解析为目录物品,并利用复用共享编码器状态的物品级排序模块对候选进行排序。为在不向服务路径添加昂贵的第二模型的情况下迁移细粒度的生产排序偏好,我们将仅用于训练的高容量教师排序器蒸馏到排序模块中。Gryphon-v2采用Rollout蒸馏进行训练:教师分数是唯一的排序监督信号,且通过两种互补的候选分布收集。当前解码器的Rollout使排序模块能够接触到服务时所用同一生成机制产生的候选,而日志展示则覆盖了用户实际看到的物品。在Yandex Music的大规模推荐场景中进行的在线A/B测试显示,单个Gryphon-v2模型替代了包含15个以上候选生成器、预排序和最终排序的生产级联。在与生产级联相当的服务延迟下,该部署使活跃用户数量增加了1.41%。这些结果证明,带有从教师排序器蒸馏而来的排序模块的生成式检索器作为生产级联的端到端替代方案具有实际可行性。

英文摘要

Industrial recommender systems are commonly deployed as multi-stage cascades with separate candidate generators, pre-rankers, and final rankers. Although effective, these cascades require repeated user-history processing, complex feature pipelines, and multiple serving stages. Semantic-ID-based generative retrieval offers a path toward simpler end-to-end systems, but next-item prediction alone does not capture the fine-grained preferences encoded by production ranking objectives. We present Gryphon-v2, a unified generate-and-rank architecture for end-to-end recommendation. The model encodes a user history once, generates Semantic-ID candidates with an autoregressive decoder, resolves them to catalogue items, and ranks them with an item-level Ranking Module that reuses the shared encoder states. To transfer fine-grained production ranking preferences without adding an expensive second model to the serving path, we distill a high-capacity, training-only Teacher Ranker into the Ranking Module. Gryphon-v2 is trained with Rollout Distillation: teacher scores are the only ranking supervision, and they are collected over two complementary candidate distributions. Rollouts from the current decoder expose the Ranking Module to candidates produced by the same generation mechanism used at serving time, while logged impressions cover items users were actually shown. In an online A/B experiment on a large-scale recommendation surface at Yandex Music, a single Gryphon-v2 model replaces a production cascade comprising more than 15 candidate generators, pre-ranking, and final ranking. The deployment increases the number of active users by 1.41% at serving latency comparable to the production cascade. These results support the practical viability of a generative retriever with a Ranking Module distilled from the Teacher Ranker as an end-to-end alternative to a production cascade.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑