TransRetrieval:面向工业推荐的基于Transformer的检索框架的规模化
TransRetrieval: Scaling Up Transformer-Based Retrieval for Industrial Recommendation
浏览论文内容
中文总结 AI 辅助
针对工业推荐检索的特征异质性问题,提出TransRetrieval框架,通过加权平均聚合、目标token压缩等技术实现高效缩放,在数据集和基准测试中取得召回率提升,线上测试也实现收益增长。
中文摘要 AI 辅助
将缩放定律应用于推荐检索受到特征异质性的阻碍:直接堆叠Transformer层会产生收益递减,因为异构字段会导致严重的token范数发散。我们提出了TransRetrieval,一种基于Transformer的检索框架,可随计算预算和跨域数据进行缩放。关键的实现手段是:(1)加权平均聚合,恢复Transformer所依赖的同构token假设;在此基础上,我们引入(2)目标token压缩,将每个候选样本的浮点运算量(FLOPs)降低85%,同时保留交叉注意力的表达能力;以及(3)位置式域嵌入,以可忽略的额外成本统一多个域,将跨域数据转化为可缩放的资产。在包含400亿次交互的工业数据集和公开的KuaiRand基准测试中,当每个目标的计算量从0.1 MFLOPs提升至2 MFLOPs时,Recall@2000分别提升19.3和22.2个百分点,证实了稳健的对数线性缩放特性。在线A/B测试中,在与生产基线相同的端到端延迟约束下,TransRetrieval使平台收益提升2.53%。
英文摘要
Applying scaling laws to recommendation retrieval is hindered by feature heterogeneity: naively stacking Transformer layers yields diminishing returns because heterogeneous fields produce severe token-norm divergence. We present TransRetrieval, a Transformer-based retrieval framework that scales with both computational budget and cross-domain data. The key enabler is (1) weighted average aggregation, which restores the homogeneous-token assumption Transformers rely on. Building on this, we introduce (2) target token compression that cuts per-candidate FLOPs by 85% while preserving cross-attention expressiveness, and (3) position-style domain embeddings that unify multiple domains at negligible additional cost, turning cross-domain data into a scaling asset. On a 40-billion-interaction industrial dataset and the public KuaiRand benchmark, scaling compute from 0.1 to 2 MFLOPs per target yields +19.3/+22.2 pt Recall@2000, confirming robust log-linear scaling. In online A/B tests, TransRetrieval lifts platform revenue by 2.53% under the same end-to-end latency constraint as the production baseline.