arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SequenceO1:基于低秩缓存的端到端超长(100K)序列推荐建模

SequenceO1: End-to-End Ultra-Long (100K) Sequence Modeling in Recommendation with Low-Rank Caching

Lin Guan, Jia-Qi Yang, Zhishan Zhao, Jiaqi Huang, Hangyu Wang, Longbin Li, Beichuan Zhang, Haonan Jiang, Jinan Ni, Xiangyu Fan, Xiaowen Li, Ziyao Ren, Yuhang Qi, Xiaolong Zhu, Xuanyuan Luo, Qiwei Chen, Yi Cheng, Lele Yu

arXiv 2609.08443首次发表:更新:

发表机构

ByteDance(字节跳动)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

SequenceO1提出端到端超长序列推荐框架,采用草图注意力压缩与低秩缓存,在抖音100K历史下实现高效训练推理并取得显著收益。

AI 中文摘要

建模长期用户行为是序列推荐和十亿级工业推荐系统的核心,然而生产级排序模型在严格的延迟、内存、通信和训练吞吐量约束下运行。在100K规模下,挑战不仅限于注意力复杂度:原始序列特征必须在训练和在线服务期间被存储、传输并反复处理。现有的基于历史截断、多阶段行为检索、压缩的终身历史或短训练/长推断外推的方法,要么削弱了端到端优化,要么保留了显著的与长度相关的成本。我们提出了SequenceO1,一个用于超长用户行为序列建模的端到端框架,已在抖音上以全流量部署,历史交互长度可达100K。SequenceO1遵循先压缩后推理的设计。其草图注意力(Sketch Attention, SA)使用可学习的原型和原型级归一化,将原始历史压缩为固定大小、与目标无关的用户表示。随后,目标条件的堆叠式目标到历史交叉注意力(Stacked Target-to-History Cross Attention, STCA)对互补的时间尺度进行建模:最近10K后缀用于短期兴趣,紧凑草图用于长期偏好。为使训练和推理切实可行,SequenceO1结合了低秩用户表示缓存、多请求用户级批处理、流水线提升和融合的FlashSA内核,以在目标、训练实例和连续请求之间分摊特征存储、通信和计算。生产实验显示了一致的离线和在线收益,而紧凑的缓存草图保留了直接将端到端序列排序扩展到100K的大部分好处。这些结果为高效注意力、序列压缩以及可扩展的长序列和长上下文推荐系统提供了一种实用的模型-系统方法。

英文摘要

Modeling long-term user behavior is central to sequential recommendation and billion-scale industrial recommender systems, yet production ranking models operate under strict latency, memory, communication, and training-throughput constraints. At the 100K scale, the challenge extends beyond attention complexity: raw sequence features must be stored, transferred, and repeatedly processed during training and online serving. Existing approaches based on history truncation, multi-stage behavior retrieval, compressed lifelong histories, or train-short/infer-long extrapolation either weaken end-to-end optimization or retain substantial length-dependent cost. We present SequenceO1, an end-to-end framework for ultra-long user behavior sequence modeling, deployed at full traffic on Douyin with histories of up to 100K interactions. SequenceO1 follows a compress-then-reason design. Its Sketch Attention (SA) uses learnable prototypes and prototype-wise normalization to compress the raw history into a fixed-size, target-agnostic user representation. Target-conditioned Stacked Target-to-History Cross Attention (STCA) then models complementary time scales: a recent 10K suffix for short-term interests and the compact sketch for long-term preferences. To make training and inference practical, SequenceO1 combines low-rank user representation caching, multi-request user-level batching, pipeline lift, and a fused FlashSA kernel to amortize feature storage, communication, and computation across targets, training instances, and consecutive requests. Production experiments show consistent offline and online gains, while the compact cached sketch retains most of the benefit of directly scaling end-to-end sequence ranking to 100K. These results provide a practical model-system approach to efficient attention, sequence compression, and scalable long-sequence and long-context recommendation systems.

CommentsRecSys'26 Industry Track, accepted as a long oral presentation. Production deployment on Douyin. Topics: industrial recommender systems, sequential recommendation, ultra-long user behavior sequence modeling, long-term user modeling, end-to-end ranking, CTR prediction, efficient attention, sequence compression, user representation caching, and large-scale recommendation systems

DOI:10.1145/3773078.3831864

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑