面向生成式推荐的难度感知语义-ID优化
Difficulty-Aware Semantic-ID Optimization for Generative Recommendation
浏览论文内容
中文总结 AI 辅助
针对生成式推荐中标准GRPO与树结构任务适配差的问题,提出DASO方法,通过前缀匹配深度分配部署组,在公开及内部推荐任务上均取得优于对比方法的性能。
中文摘要 AI 辅助
基于语义-ID(Semantic-ID)的生成式推荐将检索与排序建模为对层级化物品标识符的自回归生成过程。常见范式为SFT(监督微调)后接GRPO(生成式偏好优化),但标准GRPO与该树结构任务适配性较差。在冻结SFT检查点的情况下,对于许多提示,50束约束排序的前16个候选中不存在精确目标,更难的场景下这些候选中无一个进入目标SID分支。这种提示级诊断引出训练问题:当在线策略GRPO组同样缺失目标时,即使部分候选遵循目标路径的部分片段,物品级奖励也可能产生微弱或退化的奖励变化。我们提出难度感知语义-ID优化(DASO),一种感知树结构的后训练方法,将该失败模式视为在线部署分配问题。DASO不使用固定难度桶或均匀注入真实完成内容,而是通过前缀匹配深度刻画当前每个部署组,定位候选偏离目标路径的瓶颈SID层级,并将组的限定部分重新分配给前缀引导的完成内容,同时保留原始部署内容用于对比。SID前缀奖励提供分级信用,而辅助SFT锚点缓解SFT检查点已解决示例的退化问题。在公开基准上,DASO在12个指标中的11个优于MiniOneRec风格的GRPO,且在12个指标中的9个取得最佳结果;在内部推荐任务上,其还提升了多数层级召回指标。
英文摘要
Semantic-ID-based generative recommendation casts retrieval and ranking as autoregressive generation over hierarchical item identifiers. A common recipe is SFT followed by GRPO, yet vanilla GRPO is poorly matched to this tree-structured task. Under the frozen SFT checkpoint, the exact target is absent from the first 16 candidates of the 50-beam constrained ranking for many prompts, and in harder cases none of these candidates enters the target SID branch. This prompt-level diagnostic motivates a training concern: when on-policy GRPO groups are similarly target-missing, item-level rewards may produce weak or degenerate reward variation even if some candidates follow part of the target path. We propose Difficulty-Aware Semantic-ID Optimization (DASO), a tree-aware post-training method that addresses this failure mode as an online rollout-allocation problem. Instead of using fixed difficulty buckets or uniformly injecting ground-truth completions, DASO profiles each current rollout group by prefix-match depth, locates the bottleneck SID levels where candidates leave the target path, and reallocates a bounded portion of the group to prefix-guided completions while retaining raw rollouts for contrast. A SID-prefix reward provides graded credit, while an auxiliary SFT anchor mitigates regression on examples already solved by the SFT checkpoint. On the public benchmarks, DASO improves over MiniOneRec-style GRPO on 11 of 12 metrics and achieves the best result on 9 of 12 metrics; it also improves most level-wise recall metrics on the internal recommendation task.
发表机构
- Meta
- The Pennsylvania State University(宾夕法尼亚州立大学)
机构由 AI 辅助整理,请以论文原文为准。