arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.17092cs.IR

不确定性作为补救措施:减轻短视频多目标集成排序中的满意度标签偏差

Uncertainty as Remedy: Mitigating Satisfaction Label Bias in Short Video Multi-Objective Ensemble Ranking

Zonghe Shao, Tiantian He, Xiaoxiao Xu, Jiaqi Yu, Minzhi Xie, Jinfang Gu, Yongqi Liu, Kaiqiao Zhan, Kun Gai

AI总结:

研究短视频推荐中多目标集成排序模型因用户行为信号带来的满意度标签偏差问题,提出UAME框架,将模型预测表示为高斯评分变量,设计概率成对排序损失和加权方案减轻偏差,实验证明其能改进现有范式并与用户满意度更好对齐。

AI中文摘要:

短视频推荐的核心目标是对用户对推荐视频的真实满意度进行建模。作为主导的工业框架,端到端多目标集成排序模型通常使用多维密集用户行为信号(如点击和观看时间)进行训练。然而,这些行为信号是部分的、碎片化的,且常常相互冲突,在满意度建模中引入了不确定性和标签偏差。传统确定性模型忽略了这种不确定性,加剧了满意度标签偏差并导致次优模型收敛。现有不确定性感知方法大多将不确定性用于事后排序调整,而非作为减轻核心优化流程中固有偏差的补救措施。本文提出了UAME,一种用于短视频推荐的不确定性感知端到端多目标集成排序框架。UAME将模型预测表示为高斯评分变量,均值表示预测满意度得分,方差量化与该得分相关的预测不确定性。我们进一步设计了概率成对排序损失,并构建了不确定性感知样本级加权方案以减轻偏差。理论分析表明该加权方案有助于减轻满意度标签偏差。在大规模工业短视频平台上进行的广泛离线和在线实验表明,UAME持续改进了两种先进范式EMER和EASQ,并更好地与基于问卷的用户满意度保持一致。UAME已部署在我们的生产短视频推荐系统中,并持续带来稳定且具有统计学意义的数据提升。

英文摘要:

The core objective of short video recommendation is to model users' unobservable true satisfaction with recommended videos. As the dominant industrial framework, end-to-end multi-objective ensemble ranking models are typically trained with multi-dimensional dense user behavioral signals, such as clicks and watch time. However, these behavioral signals are partial, fragmented, and often mutually conflicting user satisfaction proxies, introducing uncertainty and label bias into satisfaction modeling. Conventional deterministic models overlook this uncertainty, which exacerbates satisfaction label bias and results in suboptimal model convergence. Meanwhile, existing uncertainty-aware methods mostly employ uncertainty for post-hoc ranking adjustments rather than leveraging it as a remedy to mitigate the inherent bias within the core optimization pipeline. This paper proposes UAME, an Uncertainty-Aware end-to-end Multi-objective Ensemble ranking framework for short video recommendation. UAME represents the model's prediction as a Gaussian scoring variable, where the mean denotes the predicted satisfaction score and the variance quantifies predictive uncertainty associated with this score. We further design a probabilistic pairwise ranking loss, and construct an uncertainty-aware sample-level weighting scheme to mitigate the bias. We further provide theoretical analysis suggesting that the weighting scheme helps mitigate satisfaction label bias. Extensive offline and online experiments on a large-scale industrial short video platform demonstrate that UAME consistently improves two state-of-the-art paradigms, EMER and EASQ, and better aligns with questionnaire-based user satisfaction. UAME has been deployed in our production short-video recommendation system and continues to deliver stable, statistically significant gains.

↑