arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

从兴趣到语义ID:面向生成式推荐的检索归因信用分配

From Interests to Semantic IDs: Retrieval-Grounded Credit Assignment for Generative Recommendation

Mengdan Zhu, Yufan Zhao, Yao Zhao, Sophie Di, Tao Di, Yulan Yan, Sridhar Iyer, Liang Zhao

arXiv 2609.29983首次发表:更新:

发表机构

Emory University; Microsoft; Cornell University(埃默里大学; 微软; 康奈尔大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对生成式推荐中语义ID奖励稀疏导致的信用分配缺口,提出检索归因的查询归因方法,将轨迹分解为历史、兴趣假设和最终ID,用冻结检索器逐假设验证并分配跨度级信用,在三个数据集上提升推荐性能。

AI 中文摘要

语义ID(SIDs)将每个目录项编码为短标记序列,使生成式推荐器能够自回归地预测下一项。推理增强变体作为一种日益常见的扩展,首先生成文本轨迹,然后通过束搜索解码下一项的语义ID。此类推荐器通常使用组相对策略优化进行训练,并采用精确匹配的语义ID奖励,该奖励在大型目录中较为稀疏。由此产生两种失败模式:当组内所有轨迹均未命中目标时,该组产生零优势且无学习信号;共享相同语义ID奖励的轨迹获得相同的优势,无论其轨迹差异多大。在这两种情况下,奖励仅反映解码出的语义ID,而从不反映产生该ID的推理过程。这造成了信用分配缺口。我们通过检索归因的查询归因来解决这一缺口。每条轨迹被结构化为历史摘要、一组兴趣假设和最终语义ID。一个冻结的检索器将每个假设作为目录查询执行,从而使每个假设可独立验证,而非仅通过最终语义ID进行评判。当任一查询在top-K内检索到目标时,该轨迹获得奖励,且逐查询命中指示器将奖励定位到各个假设。因此,信用在跨度层面进行分配:只有单独命中的假设获得正检索优势,而检索通道从不更新最终语义ID跨度。共享语义ID奖励的轨迹因此可接受不同的更新。在三个Amazon Reviews数据集上的实验中,该方法在语义ID推荐方面带来了一致的改进。在Video Games数据集上,一项预言机分析进一步揭示了兴趣条件化语义ID解码的潜力:在生成的兴趣中选择目标相关查询可同时提升召回率和排序质量。

英文摘要

Semantic IDs (SIDs) encode each catalog item as a short token sequence, enabling generative recommenders to predict the next item autoregressively. Reasoning-enhanced variants, an increasingly common extension, first generate a textual trace and then decode a next-item SID by beam search. Such recommenders are commonly trained with group-relative policy optimization under an exact-match SID reward, which is sparse in large catalogs. Two failure modes follow. When all rollouts in a group miss the target, the group yields zero advantage and no learning signal. Rollouts sharing the same SID reward receive identical advantages, however much their traces differ. In both cases the reward reflects only the decoded SID, never the reasoning that produced it. This creates a credit-assignment gap. We address this gap with retrieval-grounded query attribution. Each trace is structured into a history summary, a set of interest hypotheses, and a final SID. A frozen retriever executes every hypothesis as a catalog query, so that each hypothesis becomes independently verifiable rather than judged only through the final SID. A rollout is rewarded when any of its queries retrieves the target within the \mbox{top-$K$}, and per-query hit indicators localize that reward to individual hypotheses. Credit is thus assigned at the span level: only hypotheses that individually hit receive positive retrieval advantage, while the retrieval channel never updates the final SID span. Rollouts that share a SID reward can therefore receive different updates. Across experiments on three Amazon Reviews datasets, this yields consistent improvements in SID recommendation. On Video Games, an oracle analysis further reveals the potential of interest-conditioned SID decoding: selecting the target-relevant query among generated interests improves both recall and ranking.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑