arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

DrEM:面向视频推荐中含噪用户偏好预测的双侧鲁棒集成排序

DrEM: Dual-Side Robust Ensemble Ranking from Noisy User Preference Predictions in Video Recommendation

Canwei Huang, Tiantian He, Xiaoxiao Xu, Jun Zhang, Ziran Deng, Weike Pan, Chunjie Chen, Kaiqiao Zhan

arXiv 2608.12778首次发表:更新:

AI 中文总结

针对视频推荐中上游多任务模型输出的用户偏好预测含噪导致的排序不稳定问题,提出双侧鲁棒集成排序框架DrEM,通过风险去噪鲁棒损失与排序一致性正则化提升性能,经实验验证有效。

AI 中文摘要

工业视频推荐系统通常采用多阶段架构,在集成排序阶段,来自上游多任务模型的多维用户偏好预测(pxtrs)会被融合为统一排序分数以反映用户满意度。由于用户真实满意度难以直接观测,集成排序模型通常将pxtrs同时作为输入特征和构造代理偏好的来源。但作为上游预测模型的输出,pxtrs不可避免地包含预测噪声,该噪声会通过两个侧面向下游学习传播:在监督侧,含噪pxtrs可能翻转代理偏好并引入错误梯度;在特征侧,pxtr噪声可能通过模型输入传播,导致排序分数不稳定。现有集成排序方法通常将pxtrs视为可靠信号,忽略此类预测噪声。为解决该问题,本文提出DrEM,一种双侧鲁棒集成排序框架。DrEM引入风险去噪鲁棒损失,通过估计的偏好翻转概率修正经验风险;同时,它从预测噪声分布中采样扰动,引入偏好保持排序一致性正则化项以提升特征侧输出稳定性。理论上,本文得到了预测噪声的近似分布,并证明在偏好翻转概率估计误差下,该鲁棒损失仍具优越性。大量离线实验和大规模在线A/B测试验证了DrEM的有效性与鲁棒性。

英文摘要

Industrial video recommendation systems typically adopt a multi-stage architecture. At the ensemble ranking stage, multi-dimensional user preference predictions (pxtrs) from an upstream multi-task model are fused into a unified ranking score to reflect user satisfaction. Since users' true satisfaction is difficult to observe directly, ensemble ranking models commonly use pxtrs both as input features and as a source for constructing proxy preferences. However, as outputs of an upstream prediction model, pxtrs inevitably contain prediction noise, which propagates to downstream learning across two sides. On the supervision side, noisy pxtrs may flip proxy preferences and introduce erroneous gradients. On the feature side, pxtr noise may propagate through model inputs and destabilize ranking scores. Existing ensemble ranking methods typically treat pxtrs as reliable signals and overlook such prediction noise. To address this, we propose DrEM, a dual-side robust ensemble ranking framework. Our DrEM introduces a risk-denoising robust loss that corrects the empirical risk using estimated preference flip probability. Meanwhile, it samples perturbations from the distribution of prediction noise and introduces a preference-preserving ranking consistency regularizer to improve feature-side output stability. Theoretically, we obtain an approximate distribution of the prediction noise and prove that the robust loss remains superior under flip probability estimation error. Extensive offline experiments and large-scale online A/B tests demonstrate the effectiveness and robustness of our DrEM.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑