发表机构
Xiaohongshu; Peking University; Beijing Jiaotong University(小红书; 北京大学; 北京交通大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文针对基于LLM的推荐系统提出分层潜在推理框架HiLaR,通过层感知强化优化提升推荐性能,在四个Amazon基准数据集上优于多种主流基线。
AI 中文摘要
大语言模型(LLM)凭借其语义理解与上下文建模能力,在推荐领域展现出强大潜力,近期研究引入推理机制以优化用户偏好建模。然而,显式自然语言推理会产生巨大推理开销,现有潜在推理方法主要聚焦于中间状态的生成或验证,未能充分刻画各层的偏好角色与贡献。本文提出HiLaR,这是一种面向基于LLM的推荐系统的分层潜在推理框架,具备层感知强化优化能力。HiLaR构建时间引导的分层用户偏好表示,将其与多个LLM潜在推理状态对齐,并将推理过程从宽泛偏好组织为细粒度当前意图。为进一步优化推理轨迹,HiLaR结合最终推荐反馈与各状态边际目标似然增益导出的层感知过程奖励。在四个Amazon基准数据集上的实验表明,HiLaR总体优于强大的序列式、生成式及基于LLM的推荐基线; ablation与敏感性分析进一步验证了分层表示学习、潜在对齐及过程级优化的贡献。代码可在指定URL获取。
英文摘要
Large Language Models (LLMs) have shown strong potential for recommendation by leveraging their semantic understanding and contextual modeling capabilities. Recent studies further introduce reasoning mechanisms to improve user preference modeling. However, explicit natural-language reasoning incurs substantial inference overhead, whereas existing latent reasoning methods mainly focus on generating or verifying intermediate states, leaving their layer-wise preference roles and contributions insufficiently characterized. We propose HiLaR, a Hierarchical Latent Reasoning framework with layer-aware reinforcement optimization for LLM-based recommendation. HiLaR constructs temporal-guided hierarchical user preference representations, aligns them with multiple LLM latent reasoning states, and organizes the reasoning process from broad preferences to fine-grained current intents. To further optimize the reasoning trajectory, HiLaR combines final recommendation feedback with layer-aware process rewards derived from the marginal target-likelihood gain of each state. Experiments on four Amazon benchmark datasets show that HiLaR generally outperforms strong sequential, generative, and LLM-based recommendation baselines. Ablation and sensitivity analyses further verify the contribution of hierarchical representation learning, latent alignment, and process-level optimization. Our code is available in https://github.com/hupeiyu21/HiLaR.