发表机构
Netflix(网飞公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出用于Netflix大规模推荐解释评估的LLM评判模型生命周期框架,经五周A/B测试证实,其对齐的解释可提升会员对新颖内容的观看及浏览到播放会话量。
AI 中文摘要
LLM评判模型(LLM-as-a-Judge)是一种利用大语言模型评估其他AI应用或模型生成的自然语言的方法,现已成为加速和扩展成本高昂的人工评估的标准、可扩展方式。然而,大多数研究将评判模型视为静态产物,仅在构建时或针对固定基准进行一次评估。相反,我们认为,在生产系统中运行的LLM评判模型更应被理解为具有生命周期:它必须构建、训练、部署,并随着周围数据的演变而持续维护,每个阶段都带来独特的技术和运营挑战。我们提出了Netflix用于评估面向用户的推荐解释的LLM评判模型的生命周期,该流程每周生成并由评判模型评估数十万条不同的节目级解释,在移动端为数百万会员提供服务。我们的框架包含四个阶段:(I)诞生阶段,定义多个评估标准并构建带有人工标签和理由的精选基准数据集;(II)训练阶段,通过推理对齐 rubric 调优(Reasoning-Aligned Rubric Tuning,RART)优化评判模型的 rubric,这是一种使用基于推理输出的元评判模型作为学习信号的 rubric 调优流程;(III)部署阶段,一个评判模型承担两个生产角色:质量把关和反思性生成;(IV)监控阶段,这是一个持续的人在环对齐流程,用于检测漂移并在人工审核闸门前触发重新调优。我们报告了发布后五周针对数千万会员的A/B测试结果,其中经评判模型对齐的解释使会员的观看转向了新颖内容(此前未观看过的内容),并相对于无解释对照组增加了成功的浏览到播放会话,且未出现与质量相关的下架情况。
英文摘要
LLM-as-a-Judge, which leverages a large language model to evaluate natural language generated by another AI application or model, has become a standard, scalable approach for accelerating and extending costly human evaluation. Yet most work treats a judge as a static artifact, evaluating it once at construction or against a fixed benchmark. We argue instead that an LLM judge operating in a deployed system is better understood as having a lifecycle. It must be built, trained, deployed, and continuously maintained as the surrounding data evolves, and each phase poses distinct technical and operational challenges. We present such a lifecycle for the LLM judges that evaluate recommendation explanations at Netflix. Everything we report comes out of a series of controlled online member-facing experiments, in which our pipeline generated and the judges assessed hundreds of thousands of distinct show-level explanations per week across a changing catalog. Our framework has four phases. (I) Birth defines the evaluation criteria and builds curated benchmark datasets with human labels and rationales. (II) Training refines the judges' rubrics via Reasoning-Aligned Rubric Tuning (RART), which uses a meta-judge over reasoning output as the learning signal. (III) Deployment puts one judge in two online roles, quality gating and reflective generation. (IV) Monitoring runs a continuous Human-in-the-Loop (HITL) alignment process that detects drift and triggers re-tuning behind a human review gate. We report results from a five-week online A/B test over tens of millions of members on the Netflix mobile app, in which judge-aligned explanations shifted member viewing toward novel content (previously unwatched) and increased successful browse-to-play sessions relative to a no-explanation control, with no quality-related escalations.