AI 中文总结
针对流式推荐的行为分布漂移问题,提出知识-几何解耦(KGD)方法,通过行为多令牌预测和解耦参数所有权实现高效预训练迁移,在8个基准及Shopee生产场景中均取得显著提升。
AI 中文摘要
工业推荐系统越来越多地采用预训练后迁移的范式,但行为分布漂移引发了两个问题:应从行为序列中学习什么,以及在预训练模型持续刷新时如何迁移所学知识。为解决这些问题,我们提出了知识-几何解耦(Knowledge-Geometry Decoupling,KGD)。对于应学习什么,传统的下一个令牌预测将邻接关系视为依赖关系,可能会编码跨不相关会话的虚假转移。我们引入了行为多令牌预测(Behavioral Multi-Token Prediction,BMTP),仅保留协同或语义相关的未来项作为监督,从而产生更清晰、更具可迁移性的行为知识。对于如何迁移,预训练知识和特定任务的几何结构对共享参数施加了冲突的优化需求。为处理这一点,KGD将它们分配到单独的参数集中:可刷新的编码器拥有行为知识,而任务学习者通过只读交叉注意力读取上下文编码的状态,并通过与预训练嵌入正交的锚定校准残差(Anchored Calibration Residual,ACR)写入特定任务的几何结构。这种解耦的所有权使得能够持续刷新知识,而不会受到任务梯度的干扰或使下游适应失效。KGD在8个公共基准上比强大的预训练迁移基线提高了4-12%,并且在90天的生产流中保持了其优势,而基线则没有任何增益。KGD已在Shopee全面部署,在Shopee首页搜索的实时A/B测试中,它将每用户GMV提高了1.75%,广告收入提高了1.53%,证明了其高实用价值。我们在该httpsURL提供了KGD的核心实现。
英文摘要
Industrial recommenders increasingly adopt the pretrain-then-transfer paradigm, yet behavioral distribution drift raises two questions: what to learn from behavior sequences, and how to transfer the learned knowledge while the pretrained model is continually refreshed. To resolve them, we propose Knowledge-Geometry Decoupling (KGD). For what to learn, conventional next-token prediction treats adjacency as dependency and may encode spurious transitions across unrelated sessions. We introduce Behavioral Multi-Token Prediction (BMTP) to retain only collaboratively or semantically related future items as supervision, yielding cleaner and more transferable behavioral knowledge. For how to transfer, pretrained knowledge and task-specific geometry impose conflicting optimization demands on shared parameters. To handle it, KGD assigns them to separate parameter sets: a refreshable encoder owns behavioral knowledge, while a task learner reads contextualized encoder states through read-only cross-attention and writes task-specific geometry through Anchored Calibration Residual (ACR) orthogonal to the pretrained embedding. The decoupled ownership enables continual knowledge refresh without task-gradient interference or invalidating downstream adaptation. KGD improves over strong pretrain-transfer baselines by 4-12% on eight public benchmarks and sustains its advantage over a 90-day production stream where baselines show no gains. KGD has been fully deployed in Shopee. In a live A/B test on Shopee Homepage Search, it increases GMV per user by 1.75% and advertising revenue by 1.53%, demonstrating its high practical value. We provide the core implementation of KGD at https://github.com/FuCongResearchSquad/KGD4REC.
CommentsWithdrawn due to data sharing and privacy regulations of industrial co-authors