arXivDaily arXiv每日学术速递 周一至周五更新
arXiv 2608.11980cs.IRcs.AI

HCGRec:带有语义ID的提示条件生成式推荐

Learning from Unreachable Rewards: Hint-Conditioned Reinforcement Learning for Generative Recommendation

  • Shanghai Jiao Tong University(上海交通大学)
  • Mei Tuan(美团)

机构由 AI 辅助整理,请以论文原文为准。

Kangning Zhang, Haotian Fang, Xukun Luo, Hao Yin, Yang Gao, Peng Yan, Weiwen Liu, Weinan Zhang, Yong Yu

AI总结:

该研究提出HCGRec框架,通过提示感知信用分解优化语义ID生成式推荐,解决零奖励训练问题,在序列推荐基准上大幅提升性能并降低零优势样本占比。

AI中文摘要:

语义ID生成式推荐器将每个物品表示为离散语义标记的短序列,并通过自回归生成该标记序列来预测下一个物品。这种范式为物品ID、用户历史和物品文本提供了统一的生成接口,但也在基于奖励的后训练过程中造成了结构化优化瓶颈:当早期语义标记进入物品-标记空间的错误分支时,有限的回滚组很少能达到真实物品,因此组相对优化会获得相同的零奖励,无法产生有用的优势。我们提出了HCGRec(Hint-Conditioned Generative Recommendation),这是一种语义ID生成式推荐框架,用于为此类困难训练实例恢复学习信号。该框架通过检查点回滚诊断每个实例,仅在当前生成器无法达到正确物品时提供最小的目标前缀提示。随后模型在提示的语义分支下生成未提示的后缀,将零奖励组转化为对物品-标记完成情况的信息性比较。提示还会改变标记的身份:提示的前缀标记是由专家提供的物品上下文,而未提示的后缀标记是采样的生成动作。因此,我们引入了提示感知信用分解,使用监督学习来保留提示标记的物品-语义和前缀结构对齐,并使用GRPO优化采样的后缀。在序列推荐基准上的实验表明,HCGRec显著优于监督微调与普通基于奖励的后训练,同时将零优势训练样本从超过70%降至低于20%。代码可访问于此httpsURL。

英文摘要:

Semantic-ID generative recommenders represent each item as a short sequence of discrete semantic tokens and predict the next item by autoregressively generating this token sequence. This paradigm enables a unified generation interface for item IDs, histories, and item text, but it also creates a structured optimization bottleneck during reward-based post-training: when an early semantic token enters the wrong branch of the item-token space, finite rollout groups rarely reach the ground-truth item, so group-relative optimization receives identical zero rewards and produces no useful advantage. We propose Hint-Conditioned Generative Recommendation (HCGRec), a semantic-ID generative recommendation framework that recovers learning signal for such hard training instances. HCGRec diagnoses each instance with checkpoint rollouts and supplies a minimal target-prefix hint only when the current generator cannot reach the correct item. The model then generates the unhinted suffix under the hinted semantic branch, turning zero-reward groups into informative comparisons over item-token completions. Hinting also changes token identity: hinted prefix tokens are oracle-provided item context, while unhinted suffix tokens are sampled generation actions. We therefore introduce hint-aware credit decomposition, using supervised learning to preserve item-semantic and prefix-structure alignment for hinted tokens and GRPO to optimize the sampled suffix. Experiments on sequential recommendation benchmarks show that HCGRec substantially improves over supervised fine-tuning and vanilla reward-based post-training, while reducing zero-advantage training samples from over 70% to below 20%. The code is accessible at https://github.com/WncFht/GRec.

补充信息

↑