发表机构
Alibaba International Digital Commerce Group(阿里巴巴国际数字商业集团)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对工业推荐中Transformer扩展的计算消耗和负迁移问题,提出LazFormer,采用生成式预训练、可迁移残差适配器及非对称多轮训练,提升排序性能。
AI 中文摘要
Transformer在大型语言模型中因其卓越的可扩展性而展现出令人期待的性能,已有若干研究探讨了Transformer在工业推荐中的可扩展性。这些研究通常依赖单一的排序模型从头优化稀疏和稠密参数,导致大量计算资源消耗和收敛缓慢。幸运的是,预训练模型通过为后续排序提供有利的稀疏和稠密参数初始化,为上述问题提供了有效解决方案。然而,它们仍面临两大主要限制:(1)由于预训练和排序所使用的输入特征通常不一致,直接将稠密参数从预训练迁移到排序可能导致负迁移。(2)排序过程中的多轮训练可能导致稀疏参数的过拟合,而冻结稀疏参数则限制了它们对排序目标的适应性。为此,我们提出了一种用于工业推荐的可迁移生成式预训练扩展Transformer,称为LazFormer。具体而言,我们首先提出一个生成式预训练模块,自回归地生成序列特征,为后续排序提供有利的稀疏和稠密参数初始化。为解决稠密参数的负迁移问题,我们提出一种可迁移残差适配器,以残差方式将额外的排序特定特征注入排序过程。此外,一个请求感知排序模块整合了长序列压缩、混合稀疏注意力和请求感知范式,以高效建模用户的长序列。另外,我们进一步提出一种非对称多轮训练策略,在每轮中重置稀疏参数同时持续累积稠密参数,从而缓解稀疏参数的过拟合。
英文摘要
Transformers have shown promising performance in LLMs due to their outstanding scalability, several studies have investigated the scalability of Transformers for industrial recommendation. They typically rely on a single ranking model to optimize both sparse and dense parameters from scratch, resulting in substantial computational resource consumption and slow convergence. Fortunately, the pre-training models offer an effective solution to the above issues by providing favorable initialization of both sparse and dense parameters for the subsequent ranking. However, they still face two major limitations: (1) Since the input features used in pre-training and ranking are usually inconsistent, directly transferring dense parameters from pre-training to ranking may lead to negative transfer. (2) Multi-epoch training during the ranking process may result in the overfitting of sparse parameters, while freezing the sparse parameters limits their adaptability to the ranking objectives. To this end, we propose a Scaling Transformer for Industrial Recommendation via Transferable Generative Pre-training, termed LazFormer. Specifically, we first present a generative pre-training module to autoregressively generate sequential features, providing favorable initialization of both sparse and dense parameters for the subsequent ranking. To solve the negative transfer of dense parameters, we propose a transferable residual adapter that injects additional ranking-specific features into ranking in a residual manner. Moreover, a request-aware ranking module integrates long-sequence compression, hybrid sparse attention, and a request-aware paradigm to efficiently model users' long sequences. Besides, we further propose an asymmetric multi-epoch training strategy that resets sparse parameters while continuously accumulating dense parameters across epochs, alleviating the overfitting of sparse parameters.