GenRec:Netflix 一款基于大语言模型的推荐排序器
GenRec: An LLM-Backed Recommendation Ranker at Netflix
浏览论文内容
中文总结 AI 辅助
Netflix 提出基于 LLM 的推荐排序器 GenRec,通过两阶段后训练实现,在大规模 A/B 测试中,其用更少标注数据即可取得显著指标提升,推动推荐范式转变并总结了部署经验。
中文摘要 AI 辅助
大语言模型(LLM)正通过直接以自然语言对用户、内容和上下文进行更丰富的建模,重塑推荐系统。在 Netflix,我们正通过 GenRec 探索这一方向,GenRec 是一款基于自研基础 LLM 的推荐排序器。GenRec 遵循两阶段框架:第一阶段将开源 LLM 适配 Netflix 数据,使其深入理解内容目录和会员行为,同时平衡内容理解、指令遵循等能力;第二阶段利用推荐排序专用数据、标签和奖励信号对该基础模型进行后训练,旨在使排序器与业务需求及会员长期满意度对齐。本文聚焦第二阶段,以及从具有数千个人工特征的传统判别式排序器,向基于口头化用户历史和上下文的 LLM 排序器的转变。我们阐述了输入口头化与上下文工程、后训练数据构建、奖励整合、模型架构,以及基于仅预填充推理方法的成本受限服务设计。我们报告了 GenRec 与当前生产排序器模型对比的大规模 A/B 测试结果,显示使用少得多的第二阶段标注训练示例和输入信号训练的 GenRec 模型,可在离线和在线指标上取得统计显著的提升。我们讨论了基于 LLM 的推荐器如何改变推荐范式:从特征工程转向上下文工程,从定制架构转向共享基础主干,还概述了在现实资源约束下部署此类系统的实践经验。
英文摘要
Large language models (LLMs) are reshaping recommender systems by enabling richer modeling of users, content, and context directly in natural language. At Netflix, we are exploring this direction through GenRec, an LLM-backed recommendation ranker built on top of an in-house foundational LLM. GenRec follows a two-phase framework: Phase 1 adapts an open-source LLM to Netflix data, developing deep understanding of the catalog and member behavior while balancing capabilities such as content understanding and instruction following. Phase 2 post-trains this foundation model with recommendation-ranking specific data, labels, and reward signals, aiming to align the ranker with business requirements and long-term member satisfaction. This paper focuses on Phase 2 and the transition from a traditional discriminative ranker with thousands of engineered features to an LLM-backed ranker driven by verbalized user histories and context. We describe our design for input verbalization and context engineering, post-training data construction, reward integration, model architecture, and a cost-constrained serving design based on a prefill-only inference approach. We report results from a large-scale A/B test comparing GenRec against the current production ranker model, where we show that a GenRec model trained with substantially fewer Phase-2 labeled training examples and input signals can achieve statistically significant gains in offline and online metrics. We discuss how LLM-backed recommenders could shift the recommendation paradigm: from feature engineering to context engineering, and from bespoke architectures to shared foundation backbones. We also outline practical lessons for serving such systems under real-world resource constraints.