发表机构
ByteDance(字节跳动)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究针对工业推荐排序中行为序列建模的挑战,提出推荐原生Transformer框架ReST,通过双门控注意力等设计优化信号质量,分解排序结构解决计算不对称性,经测试提升广告平台AUC与核心收入指标并已部署生产。
AI 中文摘要
缩放Transformer已在语言建模中带来巨大收益,但将其移植到生产排序中的行为序列建模颇具挑战:推荐场景存在两方面差异,一是信号质量差异,行为序列存在噪声、时间不规则且监督稀疏;二是计算不对称性差异,每个请求需在严格延迟预算下针对一个共享用户历史对大量候选进行打分。我们提出ReST,一种推荐原生Transformer缩放框架。针对信号质量,它引入带有双门控注意力、旋转位置与时间嵌入、稳定残差归一化以及仅训练辅助目标的序列编码器;针对计算不对称性,它将排序分解为重型可复用编码器与轻量交叉解码器,采用无投影KV注意力和令牌特定参数化,结合用户级共享前缀训练与共享前缀服务,实现一次计算、多次解码的排序。在工业与公开基准测试中,ReST在序列长度、深度、宽度维度上缩放更一致且准确率更高,而LLM式Transformer块已达到饱和。在生产广告平台进行的为期一周的在线A/B测试显示,在50ms P99预算内,在线AUC提升1.31%,核心收入指标提升11.93%;ReST此后已完全部署到生产环境,表明行为序列缩放仍是生产排序中一个有前景且未充分开发的方向。
英文摘要
Scaling Transformers has driven large gains in language modeling, but transplanting this to behavior-sequence modeling in production ranking is challenging: recommendation differs in signal quality, where behavior sequences are noisy, temporally irregular, and sparsely supervised, and in computation asymmetry, where each request scores many candidates against one shared user history under tight latency budgets. We propose ReST, a recommendation-native Transformer scaling framework. For signal quality, it introduces a sequence encoder with dual-gated attention, rotary positional and temporal embedding, stabilized residual normalization, and training-only auxiliary objectives. For computation asymmetry, it factorizes ranking into a heavy reusable encoder and a lightweight cross decoder with projection-free KV attention and token-specific parameterization, coupling user-level shared-prefix training with shared-prefix serving for compute-once, decode-many-times ranking. Across industrial and public benchmarks, ReST achieves higher accuracy and scales more consistently along sequence length, depth, and width, where LLM-style Transformer blocks saturate. A one-week online A/B test on a production advertising platform improves online AUC by 1.31% and lifts a core revenue metric by 11.93% within a 50 ms P99 budget; ReST has since been fully deployed in production, showing that behavior-sequence scaling remains a promising, under-exploited axis for production ranking.