超越交互容量:基于递归模型的估计器缩放用于点击率预测
Beyond Interaction Capacity: Estimator Scaling with Recursive Models for CTR Prediction
- Georgia Institute of Technology(佐治亚理工学院)
- Google Research(谷歌研究院)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文提出估计器缩放作为交互容量缩放之外的互补方向,并据此设计参数高效的递归CTR模型RECAP,通过蒸馏、指数移动平均和推理路径聚合实现多源估计器整合,在多个基准上取得最先进性能并占据有利的性能-参数帕累托前沿。
AI中文摘要:
点击率预测是推荐和广告系统中的核心任务,依赖于对稀疏类别特征间交互的建模。显式交叉网络是CTR预测的中心范式,最近的进展主要来自通过更深的交叉网络和更具表达力的交叉算子来增加单一预测器的交互容量。我们重新审视了持续增加交互容量是否仍是提升预测性能的最有效方式,并发现即使容量持续增长,其收益也会迅速出现边际递减。这激发了一个我们称之为估计器缩放的互补缩放方向,即利用额外资源整合多个相关估计器,而非仅仅扩大单一预测器。通过理论分析,我们表明估计器缩放的收益受限于各估计器之间非共享预测变异量。然而,天真地利用这种变异可能代价高昂:独立训练的模型提供了显著的估计器多样性,但要求部署成本随集成规模增长。这激发了一种参数高效的估计器缩放实现,能够在不维护多个完整模型的情况下整合来自多个估计器源的多样性。基于这一观点,我们引入了递归平均预测器(RECAP),一种参数高效的递归CTR模型,在三个层面实现估计器缩放:跨独立训练模型的蒸馏、训练轨迹上的指数移动平均,以及权重共享递归骨干内推理时路径的聚合。在多个基准上的实验在标准基准上确立了新的最先进预测性能,同时将RECAP置于有利的性能-参数帕累托前沿。
英文摘要:
Click-Through Rate prediction, a core task in recommendation and advertising systems, relies on modeling interactions among sparse categorical features. Explicit cross networks are a central paradigm for CTR prediction, and recent progress has largely come from increasing the interaction capacity of a single predictor through deeper cross networks and more expressive cross operators. We revisit whether continually increasing interaction capacity remains the most effective way to improve predictive performance, and find that its benefits quickly exhibit diminishing returns even as capacity continues to grow. This motivates a complementary scaling direction that we call estimator scaling, where additional resources are used to incorporate multiple related estimators rather than only enlarging a single predictor. Through theoretical analysis, we show that the gains from estimator scaling are governed by the amount of non-shared predictive variation available across estimators. However, exploiting this variation naively can be expensive: independently trained models provide substantial estimator diversity but require deployment cost to grow with ensemble size. This motivates a parameter-efficient realization of estimator scaling that can incorporate diversity from multiple estimator sources without maintaining multiple full models. Building on this view, we introduce RECursive Averaged Predictor (RECAP), a parameter-efficient recursive CTR model that operationalizes estimator scaling at three levels: distillation across independently trained models, exponential moving averaging over training trajectories, and aggregation over inference-time routes within a weight-shared recursive backbone. Experiments across multiple benchmarks establish new state-of-the-art predictive performance on standard benchmarks, while placing the RECAP on a favorable performance-parameter Pareto frontier.