TransX:通过行为流与服务流的交叉扩展基于Transformer的推荐系统
TransX: Scaling Transformer-based Recommendation via Behavioral and Serving Stream Crossings
AI总结:
TransX通过解耦行为流与服务流建模,结合摊销服务策略,在领英推荐系统上实现CTR升6.0%、转化增益4.4%,且服务成本降低,性能优于现有模型。
AI中文摘要:
现代工业级推荐系统(RecSys)越来越多地采用基于Transformer的序列模型,新兴范式将推荐视为对统一的整体用户序列进行的下一个token预测。然而,将长期用户行为、实时服务事件等异构数据源合并为单一整体token流,会模糊它们各自的因果作用和时间特征,导致建模效率低下,训练和服务成本升高。我们提出TransX,这是一种面向生产的编码器-解码器架构,将推荐重新定义为序列到序列的动作转换问题。TransX明确将行为流建模与服务事件建模解耦,并基于近线行为编码与实时服务表示之间的可扩展交叉注意力来进行下一个动作解码。为实现低延迟、高QPS部署,TransX与摊销服务策略协同设计,该策略结合增量行为编码和每请求键值缓存,使服务延迟对行为序列长度不敏感。在领英(LinkedIn)的推荐系统上开展的大量离线实验和大规模在线A/B测试表明,TransX始终优于当前最先进的DLRMs及序列基线模型,实现了显著的点击率(CTR)提升(+6.0%)和转化增益(+4.4%),同时服务成本与现有生产模型相当,而我们协同设计的服务策略将在线计算量降低了约80%。
英文摘要:
Modern industrial recommender systems (RecSys) increasingly adopt Transformer-based sequence models, with an emerging paradigm that frames recommendation as next-token prediction over a unified monolithic user sequence. However, collapsing heterogeneous data sources -- such as long-term user behaviors and real-time serving events -- into a single monolithic token stream that obscures their distinct causal roles and temporal characteristics, leading to inefficient modeling and elevated training and serving costs. We propose TransX, a production-oriented encoder-decoder architecture that reformulates recommendation as a sequence-to-sequence action transduction problem. TransX explicitly decouples behavior-stream modeling from serving-event modeling and conditions next-action decoding on scalable cross-attention between nearline behavior encodings and real-time serving representations. To enable low-latency, high-QPS deployment, TransX is co-designed with an amortized serving strategy that combines incremental behavior encoding with per-request key-value caching, rendering serving latency insensitive to behavior sequence length. Extensive offline experiments and large-scale online A/B tests on LinkedIn's recommender systems show that TransX consistently outperforms state-of-the-art DLRMs and sequential baselines, and delivers substantial CTR lift (+6.0%) and conversion gain (+4.4%) while maintaining serving costs comparable to existing production models where our co-designed serving strategy reduces online computation by approximately 80%.