发表机构
Meta Platforms, Inc.(Meta平台公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究如何在推荐建模中统一非序列与序列特征。核心方法是基于悟空和 HSTU 架构构建 WHALE 模型,含特定模块与融合方式,并采用协同设计技术。主要贡献是在工业数据上离线实验有收益,在线也有积极效果且已用于生产系统。
AI 中文摘要
随着可扩展性在推荐建模中变得愈发重要,近期架构沿不同路径推进了两种广泛排名信号源的建模:非序列特征(包括用户、物品、上下文和交叉特征)以及用户行为历史中的序列特征。悟空和 HSTU 分别成为这些路径的代表性可扩展主干:悟空用于高阶非序列特征交互建模,HSTU 用于长用户行为序列建模。尽管它们优势互补,但结合这两种特征建模的实用架构仍未得到充分探索。我们提出了 WHALE,一种可扩展的统一推荐架构,在悟空和 HSTU 之上联合对非序列和序列特征进行建模。每个 WHALE 层包含一个悟空模块、一个 HSTU 模块和一个基于注意力的融合模块,其中悟空派生的交互表示查询 HSTU 派生的行为表示。这种设计使两个主干在整个网络中都保持活跃,并实现渐进式的悟空 - HSTU 交换,允许高阶特征交叉从长用户历史中反复检索细粒度证据。为使 WHALE 适用于工业部署,我们引入定制的 Triton 内核和其他模型 - 系统协同设计技术来提高训练和推理效率。在大规模工业推荐数据上,WHALE 在离线实验中取得了持续的收益。此外,它在适度牺牲服务吞吐量的情况下实现了积极的在线收益。该方法已部署于生产系统。总体而言,WHALE 提供了一个在工业推荐模型中可扩展统一这两种信息源的实际示例。
英文摘要
As scalability becomes increasingly important in recommendation modeling, recent architectures have advanced the modeling of two broad sources of ranking signals along separate paths: non-sequence features, including user, item, context, and cross features; and sequence features from user behavior histories. Wukong and HSTU have emerged as representative scalable backbones for these paths: Wukong scales high-order non-sequence feature-interaction modeling, while HSTU scales long user-behavior sequence modeling. Despite their complementary strengths, practical architectures that combine these two types of feature modeling remain underexplored. We present WHALE, a scalable unified recommendation architecture that jointly models non-sequence and sequence features on top of Wukong and HSTU. Each WHALE layer contains a Wukong module, an HSTU module, and an attention-based fusion module in which Wukong-derived interaction representations query HSTU-derived behavior representations. This design keeps both backbones active throughout the network and enables progressive Wukong-HSTU exchange, allowing high-order feature crosses to repeatedly retrieve fine-grained evidence from long user histories. To make WHALE practical for industrial deployment, we introduce customized Triton kernels and other model-systems co-design techniques to improve training and inference efficiency. On large-scale industrial recommendation data, WHALE achieves consistent gains in offline experiments. Additionally, it delivers positive online gains with a modest serving-throughput trade-off. The method has been deployed in production systems. Overall, WHALE provides a practical example of how these two sources of information can be scalably unified in an industrial recommendation model.