AI 中文总结
该研究提出两阶段知识蒸馏框架 TM20K,通过教师保留全 token、学生高效合并 token 实现超长序列建模,部署于字节跳动电商广告推荐系统,提升关键业务指标且成本接近现有模型。
AI 中文摘要
受益于超长行为序列建模,现有推荐系统通过同时考虑用户的长期与短期兴趣,为用户带来了更好的体验。然而,延长序列长度会给训练效率和服务吞吐量带来沉重负担。先前的方法通常对超长序列采用基于搜索或聚类的压缩方式,以细粒度信息为代价,或依赖各种轻量目标注意力结构,无法充分提取序列特征。在本文中,我们通过完整的 Transformer 建模结合两阶段知识蒸馏框架,平衡了超长序列建模的有效性与效率。首先,教师模型和学生模型均采用全注意力机制,而非纯目标序列注意力,以实现有效的序列扩展。对于学生模型,我们提出了几种简单但合理的 token 合并方法,在保持可接受性能的同时显著压缩序列长度。随后,一次性教师模型使用全部序列 token 进行大量训练,通过知识蒸馏进一步提升学生模型的性能。所提出的名为 TM20K 的范式已成功部署在字节跳动的电商广告推荐系统中,该系统将电商序列长度扩展至 20K,在关键业务指标上实现了显著提升(例如 ADSS 提升 1.036%),同时保持训练和服务成本与在线最先进模型几乎相同(例如服务延迟仅增加 5.6%)。
英文摘要
Benefiting from ultra-long behavior sequence modeling, existing recommender systems bring users a better experience via simultaneously considering their long-term and short-term interests. Nevertheless, extended sequence lengths introduce substantial burdens on training efficiency and serving throughput. Prior approaches typically utilize search-based or cluster-based compression on ultra-long sequences at the cost of fine-grained information, or rely on various lightweight target attention structures incapable of sufficient sequential feature extraction. In this paper, we balance the effectiveness and efficiency for ultra-long sequence modeling via full transformer modeling accompanied with a two-stage knowledge distillation framework. First, both teacher and student models take the full attention mechanism rather than pure target-sequence attention for effective sequence scaling. For student models, we propose several simple yet well-motivated token merge approaches, significantly compressing the sequence length while maintaining an acceptable performance. Then, a one-time teacher is heavily trained with full sequence tokens, further boosting the performance of student models via knowledge distillation. The proposed paradigm named TM20K has been successfully deployed in ByteDance's e-commerce advertising recommender system that extends the e-commerce sequence length to 20K, delivering substantial improvements in key business metrics (e.g., ADSS +1.036\%) while keeping the training and serving cost nearly the same as the online state-of-the-art model (e.g., serving latency only +5.6\%).
CommentsByteDance 20K Ultra-long Sequence Modeling for Ad E-Commerce Recommendation