发表机构
Baidu; City University of Hong Kong; Chinese University of Hong Kong(百度; 香港城市大学; 香港中文大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
MuSeR通过分层时间压缩、解耦多查询兴趣提取和多模态语义对齐,在工业系统中实现长序列多兴趣推荐,显著提升在线指标。
AI 中文摘要
超长用户行为序列携带了关于稳定且多样化偏好的丰富信号,然而工业推荐系统通常在严格的延迟和内存预算下将历史记录截断为几百个动作,导致长期兴趣未被充分利用。用户还在新闻、问答和短视频等多模态场景中追求多种异构意图,而稀疏的ID嵌入难以单独表示这些意图。我们提出了多兴趣序列表示(MuSeR),一个构建于已部署的MGS系统之上的检索框架,它整合了三个组件:(i)分层时间压缩,以全分辨率保留近期动作,同时逐步池化较旧的片段,使得每位用户$10^{4}$-$10^{5}$次交互的历史能够适配固定的服务预算;(ii)带有正交正则化的解耦多查询兴趣提取;(iii)多模态语义对齐,利用从大型语言模型提炼的文本摘要来增强稀疏的物品ID。为了工业部署,MuSeR进一步采用异步用户表示刷新与自适应缓存,以及跨异构硬件的分层波束搜索检索。在三个公开基准和一个大规模工业数据集上,MuSeR在Recall@$K$指标上持续优于强长序列和多兴趣基线。在百度APP首页信息流、发现信息流和短视频场景的在线A/B测试中,MuSeR带来了+0.26%的日活跃用户和+0.89%的总会话时长提升(两者均具有统计显著性,p<0.05),同时降低了服务延迟和成本。我们的贡献并非提出新的建模原语,而是进行系统级整合,使长期、多兴趣和多模态建模能够在实时生产管道中联合部署,并辅以维持其运行所需的工程实践。
英文摘要
Ultra-long user behavior sequences carry rich signals of stable and diverse preferences, yet industrial recommender systems typically truncate histories to a few hundred actions under strict latency and memory budgets, leaving long-term interests under-utilized. Users also pursue multiple heterogeneous intents across modalities such as news, Q&A, and short video, which sparse ID embeddings alone struggle to represent. We present Multi-interest Sequence Representation (MuSeR), a retrieval framework built on the deployed MGS system, which integrates three components: (i) hierarchical temporal compression, which retains recent actions at full resolution while progressively pooling older segments, so that per-user histories of $10^{4}$-$10^{5}$ interactions fit within a fixed serving budget; (ii) disentangled multi-query interest extraction with orthogonality regularization; and (iii) multimodal semantic alignment, which augments sparse item IDs with textual summaries distilled from a large language model. For industrial deployment, MuSeR further adopts asynchronous user-representation refresh with adaptive caching and hierarchical beam-search retrieval across heterogeneous hardware. On three public benchmarks and a large-scale industrial dataset, MuSeR consistently improves Recall@$K$ over strong long-sequence and multi-interest baselines. In online A/B tests on Baidu APP's homepage feed, discovery feed, and short-video scenarios, MuSeR yields +0.26% daily active users and +0.89% total session duration (both statistically significant, p<0.05), alongside reduced serving latency and cost. Rather than proposing a new modeling primitive, our contribution is a system-level integration that makes long-term, multi-interest, and multimodal modeling jointly deployable in a real-time production pipeline, together with the engineering practices required to sustain it.