发表机构
Meta(Meta)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对双塔模型训练中负样本易学习的问题,提出利用大语言模型聚类生成硬负样本的实时采样方法,能有效提升模型性能并降低流行度偏差。
AI 中文摘要
双塔模型已广泛应用于大规模推荐系统,特别是在检索阶段。训练双塔模型的行业标准通常涉及批内和/或批外负采样。然而,这些方法通常产生模型能快速学习的简单负样本,不足以充分挑战模型。为解决此问题,提出一种新颖的自监督硬负采样技术,利用大语言模型(LLM)在模型训练期间从同一聚类生成硬负样本。通过利用LLM学习媒体表示,所提方法确保生成的负样本更具挑战性和信息量。该实时采样框架设计用于无缝集成到生产模型中,能够以最小的计算复杂度处理数十亿训练数据点。在公共数据集上的实验以及在大规模在线系统中的部署表明,所提负采样技术优于广泛使用的行业方法。此外,工业应用中的分析显示,该采样方法有助于打破推荐中的固有反馈循环,并显著降低流行度偏差。
英文摘要
The two-tower model is widely used in the retrieval stage of large-scale recommendation systems, where training typically relies on in-batch and/or out-of-batch negative sampling. These methods, however, tend to produce easy negatives that the model learns quickly and that provide little training signal. This paper proposes a self-supervised, cluster-based hard negative sampling technique that draws negatives from the same semantic cluster as the positive item; in our production deployment the clusters are derived from large language model (LLM) based multimodal content representations, so that intra-cluster items are genuinely similar and yield informative negatives. To make this deployable at industrial scale, we realize the technique in a real-time, end-to-end framework that maintains a live in-memory item pool and draws cluster-based negatives from it on the fly via global out-of-batch sampling (GOOBS). The framework integrates directly into production two-tower training and serving and scales to billions of training examples with minimal computational overhead. Experiments on four public datasets and a 14-day online A/B test in a large-scale production system show that the proposed technique outperforms widely used industry methods, while also helping to break recommendation feedback loops and substantially reducing popularity bias.