arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

UniRec:级联推荐系统中的跨阶段多任务融合与偏好对齐

UniRec: Cross-stage Multi-Task Fusion with Preference Alignment for Cascaded Recommender Systems

Lingyuan Kong, Jiaqi Cui, Fanjiao Zeng, Congqi Wang, Yu Li, Yuan Cheng, Jingxin Liu, Xiaoshuang Chen, Kaiqiao Zhan

arXiv 2609.11052首次发表:更新:

发表机构

Kuaishou Technology(快手科技)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对级联推荐系统跨阶段不一致问题,提出UniRec统一融合模型,通过共享嵌入、双轴偏好对齐和属性组正则化实现跨阶段联合优化,离线优于基线,在线时长提升0.616%。

AI 中文摘要

工业推荐系统采用具有不同目标、特征空间和延迟约束的级联阶段。分别优化预排序和排序会导致跨阶段不一致:上游模型可能过滤掉下游排序器偏好的物品,而独立调优的下游融合可能抵消上游改进。现有的多任务融合方法侧重于排序阶段内的多目标融合,跨阶段方法通常仅在上游排序中添加下游得分因子。跨两个阶段的融合模块的联合优化在很大程度上仍未探索。我们提出UniRec,一种统一的跨阶段推荐融合模型。首先,两个融合代理部分共享输入嵌入,并在单个计算图中训练,因此来自任一阶段的梯度通过共享表示传播并影响另一阶段。其次,我们引入双轴偏好对齐目标:垂直跨阶段一致性项将下游成对偏好转移到上游融合得分,水平紧凑聚合项将数十个异构先验信号上的成对目标重组为双向偏好证据。第三,我们发现无约束的端到端融合优化可能利用物品属性分布的不平衡,过度集中于高奖励区域而牺牲其他目标。因此,我们添加属性组相对正则化,在属性组内计算优势,并在相同组上归一化策略,从而均匀提升整个高奖励组不会产生优化增益。离线实验中,UniRec持续优于单阶段融合和跨阶段协调基线。在线A/B测试显示应用使用时长提升0.616%。UniRec已在快手平台全面部署。

英文摘要

Industrial recommender systems cascade stages with different objectives, feature spaces, and latency constraints. Optimizing pre-ranking and ranking separately induces cross-stage inconsistency: upstream models may filter out items preferred by downstream rankers, while independently tuned downstream fusion can offset upstream improvements. Most existing multi-task fusion methods target the ranking stage alone, and cross-stage methods often align with a downstream-derived score, leaving joint optimization of fusion modules across cascaded stages largely unexplored. We propose UniRec, a Unified Cross-stage Recommendation Fusion model. First, the two fusion agents partially share input embeddings in a single computation graph, allowing gradients from either stage to propagate through the shared representations. Second, a dual-axis preference alignment objective coordinates the two stages: horizontally, a compact aggregation term reorganizes dozens of pairwise objectives over heterogeneous prior signals into bidirectional preference evidence; vertically, a cross-stage consistency term transfers downstream pairwise preferences to the upstream fusion score. Third, we introduce attribute group-relative regularization, which computes relative advantages and normalizes policy updates within each attribute group, ensuring that uniformly promoting all items in a high-reward group provides no additional optimization gain. Offline experiments demonstrate UniRec consistently outperforms single-stage fusion and cross-stage coordination baselines; online A/B experiments show a 0.616% gain in app usage duration. UniRec has been fully deployed on the Kuaishou platform.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑