WhisperRec:用于高效基础推荐模型的潜在推理方法
OneLatent: Latent Reasoning for Efficient Foundation Recommendation Models
- Kuaishou Technology(快手科技)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
WhisperRec是一种高效潜在推理框架,通过将教师CoT压缩为可学习潜在标记,在快手数据集上优于显式CoT方法,SID@64提升显著且推理吞吐量更高。
AI中文摘要:
大型语言模型(LLMs)已展现出强大的推理能力,推动其被用作基础推荐模型(FRMs)的骨干。现有方法通常在“先思考再回答”范式下,通过显式思维链(CoT)增强推荐效果。然而,生成长推理过程会带来巨大的推理开销,而固定的CoT模板难以建模多样、动态且依赖上下文的用户兴趣。我们提出WhisperRec,一种用于FRMs的高效潜在推理框架。WhisperRec将教师生成的CoT压缩为可学习的潜在推理标记,实现“先潜在推理再回答”范式,在潜在空间中执行推理而无需生成冗长的推理过程。该设计保留了与决策相关的推理信息,同时避免了自回归式推理过程生成的延迟瓶颈。具体而言,它首先引入多视图自适应思维链(MV-ACoT),从用户兴趣的互补视角构建多样且高质量的监督信号;MV-ACoT还能根据每个实例调整推理复杂度,对简单案例应用轻量级分析,对具有挑战性的案例应用针对性的多因素推理。在预训练FRM的基础上,WhisperRec采用三阶段潜在推理对齐流程,逐步将教师CoT内化到潜在表示中。最后,基于课程的后训练激活潜在标记推理以用于下游推荐,同时保留标准推荐能力。在工业级快手数据集和公开的快手LLM-Rec基准上的实验表明,WhisperRec始终优于显式CoT方法和传统基线。与显式CoT的“思考”和“不思考”变体相比,WhisperRec的SID@64分别提升了17.44%和9.33%,且在线推理吞吐量提高了10倍以上。
英文摘要:
Large language models (LLMs) have demonstrated strong reasoning capabilities, motivating their use as the backbone of foundation recommendation models (FRMs). Existing methods enhance recommendations through explicit Chain-of-Thought (CoT) reasoning under a Think-then-Answer paradigm. However, explicit CoT incurs substantial inference overhead by generating lengthy reasoning traces and relies on manually designed templates that struggle to capture diverse, dynamic user interests. We propose OneLatent, an efficient latent reasoning framework that compresses explicit reasoning traces into several learnable latent tokens, enabling Latent-Reason-then-Answer inference without generating verbose traces. OneLatent first introduces Multi-View Adaptive CoT (MV-ACoT), which creates diverse, high-quality teacher-generated supervision by exploring user interests from multiple perspectives and automatically adapting reasoning complexity to each instance. Building on pretrained FRMs, it then uses a three-stage latent-token alignment paradigm to progressively internalize CoT traces into learnable latent tokens. Finally, a multistage curriculum-based post-training strategy activates latent-token reasoning for downstream recommendation tasks. Experiments on an industrial-scale Kuaishou dataset and the public Kuaishou LLM-Rec benchmark show that OneLatent consistently outperforms explicit CoT-based methods and traditional baselines. Compared with the Think and No-Think variants of FRMs, OneLatent improves SID@64 by 17.44% and 9.33%, respectively, while achieving over 17x higher online inference throughput. We further develop a production serving system for scalable, real-time FRM inference. An online A/B test in Kuaishou's local-services advertising scenario shows that deploying OneLatent with this system yields an estimated 9.6% revenue lift over strong online baselines, including OneRec and OneReason.