AI 中文总结
研究音乐推荐器中排名模型反馈循环问题,通过离策略在线A/B测试六种干预措施及组合实验,发现服务时间干预被学习循环抵消,架构去偏有成本,不确定性驱动探索干预有提升但有权衡,给出干预层及成本建议。
AI 中文摘要
音乐推荐器中持续训练的排名模型会陷入反馈循环,即先前消费的项目主导推荐,这抑制了新发行歌曲(时间新鲜感)和未收听曲库项目(新颖性)这两类不同的内容。行业从业者有多种干预措施,学术环境对此也有一定了解,但实时系统存在持续摄入内容、组件相互连接和实际限制等挑战。本文报告了在YouTube音乐主页上针对六种干预措施和跨四个概念层(服务、训练、架构、探索)的组合实验的离策略在线A/B测试结果。所有干预措施都修改排名模型或使用其分数的服务层,候选生成和其他上游组件保持不变。结果表明,持续训练系统的服务时间干预会被学习循环抵消;架构去偏虽能减少流行度主导并提高多样性,但不会带来发现且有隐藏集成成本;使用谱归一化神经高斯过程(SNGP)头的不确定性驱动探索干预能带来最大的新发行歌曲提升,但有可衡量的参与度或多样性权衡。最后给出了关于在哪个层进行干预以及每个选择的隐藏成本的建议。
英文摘要
Continuously trained ranking models in music recommenders fall into feedback loops where previously consumed items dominate recommendations. This suppresses two distinct content classes: new releases (temporal freshness) and unlistened catalog items (novelty). Industry practitioners have a wide menu of interventions available, ranging from serving-time heuristics, training-data reweighting, architectural debiasing, to uncertainty-driven exploration, each of which are well understood in academic settings. But live systems offer challenges with continuously ingested content, interconnected components, and practical limitations that counteract the findings from academic research. We report results from off-policy online A/B tests for six interventions and a combination experiment across four conceptual layers (serving, training, architecture, exploration) on the YouTube Music homepage. All interventions modify the ranking model or the serving layer that consumes its scores; candidate generation and other upstream components are held fixed. We discuss key takeaways from our results: first, serving-time interventions on continuously trained systems are neutralized by the learning loop. Second, architectural debiasing reduces popularity dominance and improves diversity but does not create discovery, while carrying hidden integration costs. Finally, uncertainty-driven exploration interventions with a Spectral-normalized Neural Gaussian Process (SNGP) head produce the largest new-release lift, though they come with a measurable engagement or diversity tradeoff. We close with recommendations on which layer to intervene at, and the hidden costs of each choice.
CommentsAccepted at 20th ACM Conference on Recommender Systems, September 27-October 02, 2026, Minneapolis, MN, USA. 8 pages