arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.28589cs.IR

OneTrans-V2:用一个Transformer统一工业推荐系统中的召回、粗排与精排

OneTrans-V2: Unifying Retrieval, Pre-rank, and Fine-rank with One Transformer in Industrial Recommender

  • ByteDance Global E-Commerce Recommendation Foundation Team(字节跳动全球电商推荐基础团队)

机构由 AI 辅助整理,请以论文原文为准。

Hannan Cao, Jun Guo, Haolei Pei, Zhaoqi Zhang, Tianyu Wang, Ziyang Wang, Youchen Sun, Yue Xue, Yucheng Mao, Lintao Yan, Yufei Feng, Shaowei Liu, Rongkun Xing, F… 展开作者

Hannan Cao, Jun Guo, Haolei Pei, Zhaoqi Zhang, Tianyu Wang, Ziyang Wang, Youchen Sun, Yue Xue, Yucheng Mao, Lintao Yan, Yufei Feng, Shaowei Liu, Rongkun Xing, Feiling Gong, Xinyu Chenli, Cong Xu, Mingge Zhang, Yunjia Zhu, Yajing Zhang, Pengfei Ren, Yue Lin

AI总结:

OneTrans-V2用一个Transformer统一工业推荐系统的召回、粗排和精排,通过联合训练、稀疏MoE和决策条件生成式召回,提升GMV 9.74%并实现3.2倍吞吐量。

AI中文摘要:

工业推荐系统通常以召回、粗排和精排的级联方式运行,但这些阶段通常作为独立模型进行训练和服务,导致用户序列的重复编码、孤立的优化以及重复的工程工作。在OneTrans模型级统一的基础上,我们提出了OneTrans-V2,一个统一整个级联流程的Transformer。它将用户行为序列编码一次作为共享上下文,同时保留各阶段特有的候选特征和计算。联合训练使三个阶段相互增强,并支持从精排到粗排的模型内知识蒸馏。我们使用稀疏混合专家(MoE)扩展共享主干,在受限激活计算下增加容量,并通过μP风格参数化稳定扩展。为整合目标特定的召回通道,我们引入了决策条件生成式召回(DCGR)。DCGR预测描述即将发生交互的决策前缀,并基于该前缀生成物品,使业务目标能够引导单一生成过程。最后,序列原生训练(SNT)围绕每个用户的终身行为序列组织训练,并在多次曝光中摊销其编码。OneTrans-V2已部署在一个大规模工业推荐系统的所有三个阶段,将商品交易总额(GMV)提升了9.74%,并通过协同设计的服务栈,在相同硬件预算下提供了其替代级联系统3.2倍的吞吐量。

英文摘要:

Industrial recommendation systems typically operate as a \emph{cascade} of retrieval, pre-rank, and fine-rank, but these stages are usually trained and served as separate models, causing repeated user-sequence encoding, isolated optimization, and duplicated engineering effort. Building on OneTrans' model-level unification, we present OneTrans-V2, one Transformer that unifies the entire cascade. It encodes the user behavior sequence once as a shared context while preserving stage-specific candidate features and computation. Joint training lets the three stages reinforce one another and enables in-model knowledge distillation from fine-rank to pre-rank. We scale the shared backbone with sparse mixture-of-experts (MoE), which increases capacity with bounded activated computation, and stabilize scaling with $μ$P-style parameterization. To consolidate objective-specific retrieval channels, we introduce Decision-Conditioned Generative Retrieval (DCGR). DCGR predicts a decision prefix describing the upcoming interaction and generates items conditioned on it, allowing business objectives to steer a single generative process. Finally, Sequence-Native Training (SNT) organizes training around each user's lifelong behavior sequence and amortizes its encoding across exposures. Deployed across all three stages of a large-scale industrial recommendation system, OneTrans-V2 improves gross merchandise value (GMV) by 9.74\% and, with a co-designed serving stack, delivers $3.2\times$ the throughput of the cascade it replaces under the same hardware budget.

↑