arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

任务不可知不再:统一排序骨干中的多任务信息流

Task-Blind No MORE: Multi-Task Information Flow in Unified Ranking Backbones

Yuchen Wang, Feng Niu, Qing Tan, Junting Lu, Baoxin Wu, Jun Gao

arXiv 2609.07273首次发表:更新:

发表机构

Hello Group; University of Science and Technology of China; Institute of Software, Chinese Academy of Sciences; University of Chinese Academy of Sciences; Beijing Information Science and Technology University(好未来; 中国科学技术大学; 中国科学院软件研究所; 中国科学院大学; 北京信息科技大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对统一排序骨干缺乏任务感知信息流的问题,提出MORE模型,在骨干内嵌入多任务信息流,通过锚定令牌实现任务协同演化,在工业数据及线上测试中全面超越基线并显著降低延迟。

AI 中文摘要

用于推荐的工业排序模型分别扩展了特征交互和序列建模;最近的架构如HyFormer和MixFormer将两者统一到一个可堆叠的骨干中。然而,现实世界的推荐系统几乎总是需要多任务学习,但现有的统一架构将多任务建模限制在浅层的骨干后塔中,使骨干缺乏任务感知的信息流。我们提出MORE(多任务协同演化排序模型),它将多任务信息流嵌入骨干内部,使任务特定信号在每一层与序列和特征表示协同演化,而非事后融合。它引入了跨骨干层持续存在的锚定令牌:共享锚定令牌编码跨任务的共性,而私有锚定令牌捕获任务特定的先验。在每个块中,锚定令牌(1)从行为序列中读取任务条件信号,(2)在任务边界掩码下与非序列特征混合,(3)通过独立分支细化每个任务的表示;随着块的堆叠,每个任务获得通过所有骨干层细化的差异化表示。在大型工业数据集上的实验表明,在可比的参数和FLOPs预算下,MORE在所有任务上始终优于基线,并且随模型规模扩展良好。在陌陌(中国领先的社交发现平台,月活跃用户数千万)上的在线A/B测试中,使用时长提升3%,互动率提升3.6%,深度聊天率提升2%。MORE已投入生产,通过请求级共享计算将评分延迟降低约30%。

英文摘要

Industrial ranking models for recommendation have scaled feature interaction and sequence modeling separately; recent architectures such as HyFormer and MixFormer unify both in a stackable backbone. Real-world recommender systems, however, nearly always require multi-task learning, yet existing unified architectures confine multi-task modeling to shallow post-backbone towers, leaving the backbone without task-aware information flow. We propose MORE (Multi-task cO-evolving Ranking modEl), which embeds multi-task information flow inside the backbone, enabling task-specific signals to co-evolve with sequence and feature representations at every layer rather than in a post-hoc fusion. It introduces Anchor Tokens that persist across backbone layers: Shared Anchors encode cross-task commonalities, while Private Anchors capture task-specific priors. In each block, Anchor Tokens (1) read task-conditioned signals from behavior sequences, (2) mix with non-sequential features under a task-boundary mask, and (3) refine per-task representations through independent branches; as blocks stack, each task obtains a differentiated representation refined through all backbone layers. Experiments on large-scale industrial datasets show that MORE consistently outperforms baselines across all tasks under comparable parameter and FLOPs budgets, and scales well with model size. Online A/B tests on Momo, a leading Chinese social discovery platform with tens of millions of monthly active users, yield 3% improvement in usage duration, 3.6% in interaction rate, and 2% in deep-chat rate. MORE is deployed in production with request-level shared computation reducing scoring latency by about 30%.

CommentsAccepted at CIKM 2026

DOI:10.1145/3799682.3840113

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑