PRIME:通过插件式残差输入条件混合专家缓解共享CTR顶层网络中的子组优化竞争
PRIME: Mitigating Subgroup Optimization Competition in Shared CTR Top Networks with Plug-in Residual Input-Conditioned Mixture of Expert
浏览论文内容
中文总结 AI 辅助
PRIME以Dense为锚点的残差输入条件混合专家,解决共享CTR顶层网络的子组优化竞争,在Avazu、Criteo等数据集上提升AUC、降低LogLoss且优于APG,参数更少、延迟更低。
中文摘要 AI 辅助
点击率(CTR)模型在特征交互设计上各有不同,但其顶层网络通常仍是一个由所有样本共享的单一多层感知机。因此,异构的用户、物品和上下文子组会更新相同的参数;弱对齐的学习信号会使聚合梯度成为相互竞争方向之间的折中。我们在Avazu数据集上用4个模型和4个语义字段研究了这种竞争。在所有架构中,语义子组的Top-NN梯度余弦相似度低于与样本量和标签比例匹配的随机组,降幅为0.23-0.37。这种竞争促使人们提出输入条件化的专家模型,但直接替换已有的Dense映射会改变其初始功能、共享模式和容量,从而模糊了增益的来源。我们提出了PRIME(插件式残差输入条件混合专家,Plug-in Residual Input-conditioned Mixture of Expert),一种以Dense为锚点的低秩残差专家混合模型。PRIME锚定原始预测,并使用零残差初始化,使其在训练开始时与Dense基线完全匹配。依赖输入的路由权重为样本特定的logit校正分配低秩专家;多袋聚合和EMA加载偏差稳定了条件估计。我们在保留的Avazu和Criteo测试集上,针对13种CTR架构和5组配对种子评估了PRIME。中位配对AUC增益分别为+0.0022和+0.0066,LogLoss分别降低0.0011和0.0081。在FiBiNET和DCNv2上,PRIME在所有10个种子级AUC比较中均优于APG,同时在两种骨干网络上使用更少的参数和更低的推理延迟。这些结果表明,保留功能的条件残差在增加输入依赖容量的同时,保留了Dense路径及其优化稳定性。代码可在this https URL获取。
英文摘要
Click-through rate (CTR) models vary in feature-interaction design, yet their top networks usually remain a single multilayer perceptron shared by all examples. Heterogeneous user, item, and context subgroups therefore update the same parameters; weakly aligned learning signals make the aggregate gradient a compromise among competing directions. We study the competition on Avazu with 4 models and 4 semantic fields. Across all architectures, semantic subgroups show lower Top-NN gradient cosine similarity than random groups matched by sample size and label ratio, with reductions of 0.23-0.37. This competition motivates input-conditioned experts, but directly replacing an established Dense mapping changes its initial function, sharing pattern, and capacity, obscuring the source of gains. We introduce PRIME (Plug-in Residual Input-conditioned Mixture of Experts), a Dense-anchored mixture of low-rank residual experts. PRIME anchors the original prediction and uses zero-residual initialization to match the Dense baseline exactly at training onset. Input-dependent routing weights low-rank experts for example-specific logit corrections; multi-bag aggregation and EMA load biases stabilize conditional estimation. We evaluate PRIME on held-out Avazu and Criteo test sets across 13 CTR architectures and five paired seeds. Median paired AUC gains are +0.0022 and +0.0066, with LogLoss reductions of 0.0011 and 0.0081, respectively. On FiBiNET and DCNv2, PRIME outperforms APG in all ten seed-level AUC comparisons while using fewer parameters and lower inference latency on both backbones. These results show that function-preserving conditional residuals add input-dependent capacity while preserving the Dense path and its optimization stability. Code is available at https://github.com/YH-learning/PRIME.
发表机构
- Ant Group(蚂蚁集团)
- Henan Polytechnic University(河南理工大学)
- Alibaba Inc.(阿里巴巴集团)
机构由 AI 辅助整理,请以论文原文为准。