arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

LaRec:释放基于大语言模型的生成式推荐中的潜在推理能力

LaRec: Unleashing LLM-based Latent Reasoning for Generative Recommendation

Yu Xia, Zihan Lin, Wei Yang, Rui Zhong, Cheng Chen, Huan Ren, Yao Hu

arXiv 2607.24617首次发表:更新:

AI 中文总结

研究针对LLMs在推荐中现有方法的问题,提出LaRec框架。核心方法是通过潜在预训练赋予潜在推理能力,用个性化强化学习调整引导遍历多样推理路径。主要贡献是在多数据集实验中,以相当效率显著超越现有基线。

AI 中文摘要

大语言模型(LLMs)在推荐方面因卓越推理能力展现出巨大潜力。现有方法主要依赖显式思维链(CoT),存在推理文本冗长、响应时间低效问题。潜在推理旨在通过在连续潜在空间内思考平衡效率,但面临缺乏细粒度监督和单一推理路径两大挑战。为此提出LaRec,它由潜在预训练和个性化强化学习调整两个核心阶段组成。潜在预训练通过步骤级对齐和过程方向对齐为潜在空间推理提供丰富监督信号,赋予LLMs潜在推理能力。个性化强化学习调整基于用户历史兴趣为每个用户构建个性化高斯混合分布,训练时从中随机采样不同推理起点,引导LLMs在潜在空间中遍历多样推理路径,有效探索用户多方面兴趣。多个数据集实验表明LaRec在效率相当的情况下显著优于现有基线。

英文摘要

Large Language Models (LLMs) have shown great promise in recommendation due to superior reasoning abilities. However, existing methods mainly rely on explicit Chain-of-Thought (CoT), resulting in verbose reasoning texts and inefficient response times. latent reasoning aims to balance efficiency by thinking within a continuous latent space, yet it faces two major challenges: (1) Lack of Fine-grained Supervision: Latent reasoning relies solely on feedback from the final labels, providing sparse supervisory signals that struggle to effectively guide the optimization of multiple hidden reasoning steps. (2) Single Reasoning Path: The deterministic nature of latent reasoning impedes the exploration of users' diverse interests and preferences, thereby limiting the recommendation capabilities of LLMs. To address these issues, we propose \textbf{$LaRec$}, an efficient generative recommendation framework designed to unleash the potential of latent reasoning in LLMs. $LaRec$ consists of two core stages: First, we design Latent Pre-training that empowers LLMs with latent reasoning capabilities by providing rich supervisory signals to the latent space reasoning via step-level alignment and process direction alignment. Second, we introduce Personalized RL-tuning. Specifically, we construct a personalized Gaussian Mixture Distribution for each user based on their historical interests. By randomly sampling distinct reasoning starting points from this distribution during training, we guide the LLMs to traverse diverse reasoning paths within the latent space, enabling efficient exploration of user's multi-faceted interests. Experiments on multiple datasets show that $LaRec$ significantly outperforms existing baselines with comparable efficiency.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑