arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

学习更好的语义ID生成式推荐推理

Learning Better Reasoning for Generative Recommendation with Semantic IDs

Mengdan Zhu, Yufan Zhao, Sophie Di, Yao Zhao, Tao Di, Yulan Yan, Sridhar Iyer, Liang Zhao

arXiv 2609.29973首次发表:更新:

发表机构

Emory University; Microsoft; Cornell University(埃默里大学; 微软; 康奈尔大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出Evo-Rec三阶段框架,通过上下文对齐、推理轨迹筛选和强化学习优化,提升基于语义ID的生成式推荐中的推理质量,在三个基准上全面超越现有方法。

AI 中文摘要

生成式推荐将物品检索重构为序列生成,使得统一模型能够直接从用户的交互历史中生成下一个物品。语义ID通过将每个物品表示为离散编码,进一步使该范式有效且可扩展,从而在语义相关的物品之间实现知识共享。近期研究在语义ID生成之前引入显式推理,帮助模型总结用户兴趣并推断可能的偏好转移。然而,推理并非固有有益:不准确或无信息的推理可能会误导后续的物品生成,最终降低推荐性能。这提出了一个核心挑战:推荐器如何选择和学习有效的推理轨迹,并逐步从自身生成中进化出更好的推理?在这项工作中,我们提出了Evo-Rec,一个三阶段框架,用于学习更好的推理并通过强化学习进一步增强。首先,我们将语义ID与其文本和行为上下文对齐,使模型能够理解和生成物品标识符。其次,我们采样多个候选推理轨迹,并保留那些能改善真实物品预测的轨迹,通过监督微调提供更强的推理初始化。第三,我们通过带有目录约束的物品生成和排序感知推荐反馈的强化学习进一步优化推理策略。在三个Amazon Review基准上的实验表明,Evo-Rec在所有评估指标上持续优于判别式、生成式和推理增强型推荐器。这些结果证明了我们的框架在基于SID的生成式推荐中学习更好推理的有效性。

英文摘要

Generative recommendation reformulates item retrieval as sequence generation, allowing a unified model to directly generate the next item from a user's interaction history. Semantic IDs further make this paradigm effective and scalable by representing each item as discrete codes, enabling knowledge sharing among semantically related items. Recent studies introduce explicit reasoning before Semantic-ID generation, helping models summarize user interests and infer possible preference transitions. However, reasoning is not inherently beneficial: Inaccurate or uninformative reasoning may mislead subsequent item generation and ultimately degrade recommendation performance. This raises a central challenge: how can a recommender select and learn effective reasoning traces and progressively evolve toward better reasoning from its own generations? In this work, we propose Evo-Rec, a three-stage framework for learning better reasoning and further enhancing it through reinforcement learning. First, we align Semantic IDs with their textual and behavioral contexts, enabling the model to understand and generate item identifiers. Second, we sample multiple candidate reasoning traces and retain those that improve the prediction of the ground-truth item, providing a stronger reasoning initialization through supervised fine-tuning. Third, we further optimize the reasoning policy through reinforcement learning with catalog-constrained item generation and ranking-aware recommendation feedback. Experiments on three Amazon Review benchmarks show that Evo-Rec consistently outperforms discriminative, generative, and reasoning-enhanced recommenders across all evaluation metrics. These results demonstrate the effectiveness of our framework in learning better reasoning for SID-based generative recommendation.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑