arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

教师保留全部 token,学生高效合并:用于广告推荐中电商序列建模的 TM20K

Teacher Retains Full Tokens, Student Merges Efficiently: TM20K for E-Commerce Sequence Modeling in Ad Recommendation

Xinchun Li, Duoru Zheng, Wenlin Zhao, Haoran Ding, Ziyi Zhou, Jingxuan Tan, Huizhi Yang, Yuchen Jiang, Zhe Chen, Yuchao Zheng, Linlan Chen, Dongjian Wang, Dongyue Wang, Xiaosong Li, Hongyue Mao, Yaocheng Tan

arXiv 2608.07055首次发表:更新:

AI 中文总结

该研究提出两阶段知识蒸馏框架 TM20K,通过教师保留全 token、学生高效合并 token 实现超长序列建模,部署于字节跳动电商广告推荐系统,提升关键业务指标且成本接近现有模型。

AI 中文摘要

受益于超长行为序列建模,现有推荐系统通过同时考虑用户的长期与短期兴趣,为用户带来了更好的体验。然而,延长序列长度会给训练效率和服务吞吐量带来沉重负担。先前的方法通常对超长序列采用基于搜索或聚类的压缩方式,以细粒度信息为代价,或依赖各种轻量目标注意力结构,无法充分提取序列特征。在本文中,我们通过完整的 Transformer 建模结合两阶段知识蒸馏框架,平衡了超长序列建模的有效性与效率。首先,教师模型和学生模型均采用全注意力机制,而非纯目标序列注意力,以实现有效的序列扩展。对于学生模型,我们提出了几种简单但合理的 token 合并方法,在保持可接受性能的同时显著压缩序列长度。随后,一次性教师模型使用全部序列 token 进行大量训练,通过知识蒸馏进一步提升学生模型的性能。所提出的名为 TM20K 的范式已成功部署在字节跳动的电商广告推荐系统中,该系统将电商序列长度扩展至 20K,在关键业务指标上实现了显著提升(例如 ADSS 提升 1.036%),同时保持训练和服务成本与在线最先进模型几乎相同(例如服务延迟仅增加 5.6%)。

英文摘要

Benefiting from ultra-long behavior sequence modeling, existing recommender systems bring users a better experience via simultaneously considering their long-term and short-term interests. Nevertheless, extended sequence lengths introduce substantial burdens on training efficiency and serving throughput. Prior approaches typically utilize search-based or cluster-based compression on ultra-long sequences at the cost of fine-grained information, or rely on various lightweight target attention structures incapable of sufficient sequential feature extraction. In this paper, we balance the effectiveness and efficiency for ultra-long sequence modeling via full transformer modeling accompanied with a two-stage knowledge distillation framework. First, both teacher and student models take the full attention mechanism rather than pure target-sequence attention for effective sequence scaling. For student models, we propose several simple yet well-motivated token merge approaches, significantly compressing the sequence length while maintaining an acceptable performance. Then, a one-time teacher is heavily trained with full sequence tokens, further boosting the performance of student models via knowledge distillation. The proposed paradigm named TM20K has been successfully deployed in ByteDance's e-commerce advertising recommender system that extends the e-commerce sequence length to 20K, delivering substantial improvements in key business metrics (e.g., ADSS +1.036\%) while keeping the training and serving cost nearly the same as the online state-of-the-art model (e.g., serving latency only +5.6\%).

CommentsByteDance 20K Ultra-long Sequence Modeling for Ad E-Commerce Recommendation

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑