大规模物品嵌入:Yandex生态系统中基于GNN和基于ID的物品嵌入比较
Embedding Items at Scale: Comparing GNN-Based and ID-Based Item Embeddings in the Yandex Ecosystem
浏览论文内容
中文总结 AI 辅助
本文通过案例研究,在Yandex的三个产品场景中对比了基于GNN的预训练物品嵌入与端到端可训练物品嵌入,发现训练数据有限时预训练有帮助,大规模数据集上则无显著益处。
中文摘要 AI 辅助
基于Transformer的序列推荐模型会处理用户-物品交互序列,其性能高度依赖物品嵌入策略。现有方法要么使用预训练物品嵌入,要么与Transformer一起端到端学习。据我们所知,此前尚无研究在大规模工业场景下从成本和质量两方面对比这些方案。本文是一项案例研究,在Yandex的两个成熟生产推荐系统——Yandex Market和Yandex Music中,对比了预训练的工业级图神经网络(GNN)物品嵌入与端到端可训练物品嵌入;还在从Yandex Lavka生产日志中采样的低资源数据集上评估了两种方法,该数据集的数据和代码均公开用于演示。我们的结果表明,当训练数据有限时,单独的预训练阶段有帮助,但对于在大量数据集上训练的大规模模型,预训练无显著益处。
英文摘要
Transformer-based sequential recommendation models, which process sequences of user-item interactions, rely heavily on the item embedding strategy. Existing approaches either use pretrained item embeddings or learn them end-to-end with the transformer. To the best of our knowledge, no prior work has compared these options from both cost and quality perspectives in a large-scale industrial setting. This paper is a case study that compares pretrained industrial graph neural network item embeddings with end-to-end trainable item embeddings across two mature production recommendation systems at Yandex: Yandex Market and Yandex Music. We additionally evaluate both approaches on a low-resource dataset sampled from Yandex Lavka production logs, for which both the data and code are publicly available for demonstration purposes. Our results show that a separate pretraining stage helps when training data is limited, but provides no worthwhile benefit for large-scale models trained on extensive datasets.