arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

大黄蜂:用于大规模推荐系统的交错混合层构建块

Bumblebee: Interleaved Mixed-Layer Building Blocks for Large-Scale Recommendation Systems

David Bauer, Cancan Zhang, Wenshun Liu, Xiaoyi Zhang, Weijia Liu, Wanli Ma, Yue Weng, Wei Li, Rui Li, Yiyang Zhao, Tianqi Lu, Jing Qian, Huayu Li, Xiaoyi Liu, Linhong Zhu, Jerry Fu

arXiv 2607.24804首次发表:更新:

发表机构

Meta(Meta)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对推荐系统发展的两条轨道缺乏交互的问题,提出大黄蜂架构,通过交错可堆叠块设计,结合多种技术,实现模态混合与信息丰富,经实验验证该方法优于基线模型,证明交错异构单元是有前途的推荐架构范式。

AI 中文摘要

推荐系统在过去几年经历了重大变革。从传统特征交互模块到生成式下一个动作预测的转变推动了个性化内容的边界。发展主要沿着两条独立的轨道进行,一方面是序列建模方法,另一方面是特征交互方法。在本文中,我们介绍了大黄蜂,一种通过交错、可堆叠的块设计解决两个方向之间缺乏交互问题的推荐架构。每个块实现了一个微管道层,将序列个性化、基于注意力的编码和特征交叉组合成一个独立的单元。每个块产生两种特征模态的联合表示,供序列中的下一个块使用。这种机制鼓励早期和重复的模态混合,并用额外的上下文信息丰富下游特征。块之间的残差连接创建跨模态信息路径,并在不增加额外参数的情况下产生额外的预测性能。块可以通过选择性地丢弃组件进行专门化,从而在质量和吞吐量之间实现灵活的权衡。我们在大规模工业数据上评估了我们的方法,并在几个分类和回归任务中显示出比可比基线模型有一致的改进。此外,我们进行了消融研究,以确认交错组合本身是这些改进的主要驱动因素。我们的结果表明,交错异构功能单元,而不是构建深层堆栈,是下一代推荐架构的一个有前途的范式。

英文摘要

Recommendation systems have undergone significant transformations in the past years. The transition from traditional feature interaction modules to generative next-action prediction has pushed the boundaries of personalized content. Developments have largely evolved along two separate tracks. Sequence modeling approaches on the one hand and feature interaction methods on the other. In this paper, we introduce Bumblebee, a recommendation architecture that addresses the lack of interaction between the two directions through an interleaved, stackable block design. Each block implements a micro-pipeline of layers combining sequence personalization, attention-based encoding, and feature crossing into a self-contained unit. Every block produces a joint representation of both feature modalities which is consumed by the next block in the sequence. This mechanism encourages early and repeated mixture of modalities and enriches downstream features with additional contextual information. Residual connections between blocks create cross-modal information pathways and yield additional predictive performance without adding additional parameters. Blocks can be specialized by selectively dropping components, enabling flexible trade-offs between quality and throughput. We evaluate our approach on large-scale industrial data and show consistent improvements over comparable baseline models across several classification and regression tasks. Furthermore, we conduct ablation studies to confirm that the interleaved composition itself is the primary driver of these improvements. Our results suggest that interleaving heterogeneous functional units, rather than composing deep stacks, is a promising paradigm for future-generation recommendation architectures.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑