arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.00750cs.IR

用于生成式推荐的分层残差策略优化

Hierarchical Residual Policy Optimization for Generative Recommendations

Kaifeng Guo, Yiming Yang, Jingtong Gao, Guolei Zeng, Fukang Yang, Yukang Liang, Peng Jiang, Qingpeng Cai, Xiangyu Zhao

AI总结:

针对生成式推荐器后训练中标记级信用分配问题,提出HRPO框架,将项目级结果转换为标记对齐学习信号,经实验在会话效用和业务指标上取得提升。

AI中文摘要:

生成式推荐器通过自回归解码语义标识符(SIDs)来选择项目,其标记位置在项目空间上形成从粗到细的层次结构。在实践中,SID解码器通过监督式下一个标记预测进行训练,该方法模仿记录的轨迹而非直接优化下游效用,这促使人们使用结果反馈进行后训练,以引导解码向更高效用方向发展。然而,记录的反馈仅针对最终暴露的项目被观察到,导致大多数后训练方法在项目级别运行,并在所有SID标记上广播相同的终端信号,结果是标记级别的信用分配变得稀疏、高方差且依赖于层级。为此,我们提出分层残差策略优化(HRPO),这是一种后训练框架,可将项目级结果转换为密集的、标记对齐的学习信号,用于保守的逐标记改进。具体而言,HRPO首先通过基于特征的用户集群进行分组奖励平滑来估计SID前缀级效用,然后将这些效用分解为残差标记信用并将其累积为累计信用信号,最后,残差回报策略优化(RRPO)使用裁剪更新、组归一化优势和KL正则化来优化残差信用,以保持稳定性。在公共数据集和大规模商业系统中的在线A/B测试实验表明,会话级效用和关键业务指标均取得了持续提升,源代码和存档工件可用于复现。

英文摘要:

Generative recommenders select items by autoregressively decoding semantic identifiers (SIDs), whose token positions induce a coarse-to-fine hierarchy over the item space. In practice, SID decoders are trained via supervised next-token prediction, which imitates logged trajectories rather than directly optimizing downstream utility. This motivates post-training with outcome feedback to guide decoding toward higher utility. However, logged feedback is only observed for the final exposed item, causing most post-training methods to operate at the item level and broadcast the same terminal signal across all SID tokens. As a result, token-level credit assignment becomes sparse, high-variance, and layer-dependent. To this end, we propose Hierarchical Residual Policy Optimization (HRPO), a post-training framework that converts item-level outcomes into dense, token-aligned learning signals for conservative token-wise improvement. Specifically, HRPO first estimates SID prefix-level utilities via group-wise reward smoothing over feature-based user clusters. It then decomposes these utilities into residual token credits and accumulates them into credit-to-go signals. Finally, Residual-Return Policy Optimization (RRPO) optimizes the residual credits using clipped updates, group-normalized advantages, and KL regularization to preserve stability. Experiments on a public dataset and an online A/B test in a large-scale commercial system show consistent gains in session-level utility and key business metrics. Source code and the archived artifact are available for reproduction.

补充信息

↑