arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.35783cs.IRcs.LG

软课程学习用于优化新鲜与泛化推荐

Soft Curriculum Learning for Optimizing Fresh and Generalized Recommendations

  • Google(谷歌)

机构由 AI 辅助整理,请以论文原文为准。

Arnab Bhadury, Siyan Zheng, Anlan Yu, Palaksh Rungta, Jiawei Li, Changping Meng, Dapeng Hong, Chuan He, Onkar Dalal

AI总结:

针对推荐系统流行度反馈循环问题,提出可扩展的软课程学习框架,通过损失退火和图内权重调整实现动态课程节奏,在不牺牲吞吐量下提升用户满意度和新鲜内容消费。

AI中文摘要:

大规模推荐系统,尤其是短视频平台,常常受到海量流行度反馈循环的瓶颈制约。在这种环境中,当模型推荐热门物品时,它们会为“头部”物品生成大量偏斜的训练数据。这形成了一个自我强化的循环,使得检索和排序模型记住“头部”物品的模式,而牺牲了对目录中庞大“尾部”的泛化能力。尽管课程学习(CL)提供了一种强大的机制,通过系统性地让模型接触逐渐变难且频率更低的样本来打破这一反馈循环,但它在工业推荐中的采用一直受到硬件利用效率低下或需要复杂预处理技术的阻碍,因为动态数据拒绝算法往往会使硬件加速器(TPU/GPU)因主要受CPU限制而饿死。在这项工作中,我们引入了一个可扩展的软课程学习框架,专门设计用于工业规模检索和排序模型中的持续训练设置。通过利用损失退火和计算图内权重调整,而不是刚性数据过滤,我们打破了流行度反馈,并实现了动态课程节奏,同时不牺牲系统吞吐量。我们通过跨序列基础检索模型(如SASRec)、双塔检索模型和大规模持续排序模型的应用展示了实证证据。在我们短视频平台上的在线A/B测试表明,整体用户满意度和新鲜内容消费均有显著提升,且模型吞吐量没有下降。

英文摘要:

Large-scale recommender systems, particularly short-form video platforms, are often bottlenecked by massive popularity feedback loops. In such environments, as models recommend popular items, they generate an overwhelming amount of skewed training data for "head" items. This creates a self-reinforcing cycle where retrieval and ranking models memorize "head" item patterns at the expense of generalizing across the vast "tail" of the catalogue. While Curriculum Learning (CL) offers a powerful mechanism to break this feedback loop by systematically exposing models to progressively more difficult and less frequent examples, its adoption in industrial recommendation has been hampered by hardware utilization inefficiencies or the needs for complicated pre-processing techniques because dynamic data rejection algorithms tend to starve hardware accelearators (TPUs/GPUs) by becoming largely CPU-bound. In this work, we introduce a scalable Soft Curriculum Learning framework designed specifically for continuous training setups within industry-scale retrieval and ranking models. By utilizing loss annealing and in-graph weight adjustments rather than rigid data filtering, we break the popularity feedback and enable dynamic curriculum pacing without sacrificing system throughput. We demonstrate empirical evidence through applications across sequence-based retrieval models (such as SASRec), two-tower retrieval models, and large-scale continuous ranking models. Online A/B tests on our short-video platform demonstrate substantial lifts in both overall user satisfaction and fresh content consumption, all without degrading model throughput.

补充信息

↑