arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

用于生成式推荐的奖励引导解码

Reward Guided Decoding for Generative Recommendation

Ruochen Yang, Yusheng Huang, Youfeng Zheng, Shuang Wen, Liangliang Chen, Pengbo Xu, Xiaoyu Zhang, Shijun Wang, Shuang Yang, Zhaojie Liu, Lantao Hu, Wenwu Ou, Jiawei Sheng, Tingwen Liu

arXiv 2607.25344首次发表:更新:

AI 中文总结

研究针对生成式推荐解码受生成似然性主导与业务目标冲突的问题,提出奖励引导解码(RGD)框架,将价值引导解码公式化为KL正则化奖励最大化问题,通过引入奖励模型重塑搜索轨迹,经实验验证有效且已在快手平台部署。

AI 中文摘要

生成式推荐将推荐任务表述为序列到序列的自回归生成范式,但解码过程常受生成似然性主导,这可能与现实业务目标冲突。现有重排序或训练时对齐方法存在干预过晚或业务偏好变化时需昂贵模型重新训练的问题。为此,我们提出奖励引导解码(RGD),一种面向工业价值的可控解码框架。我们将价值引导解码公式化为KL正则化奖励最大化问题,得出闭式奖励引导解码分布。RGD将基础生成器视为参考策略,引入奖励模型作为测试时控制器,在不重新训练生成器的情况下重塑搜索轨迹。广泛的离线和在线实验证明了该方法在对齐个性化和业务价值方面的有效性,RGD已在快手平台部署并带来持续改进。

英文摘要

Generative recommendation formulates recommendation task into an SID sequence autoregressive generation paradigm, but the decoding process is often dominated by generation likelihood. This may conflict with real-world business objectives, where high-value candidates can receive low generation probability and be pruned early during beam search. Existing reranking or training-time alignment methods either intervene too late or require costly model retraining when business preferences change. To this end, we propose \textbf{R}eward \textbf{G}uided \textbf{D}ecoding, named \textbf{RGD}, a controllable decoding framework for industrial value-oriented generative recommendation. We formulate value-guided decoding as a KL-regularized reward maximization problem, deriving a closed-form reward guided decoding distribution that principledly combines generation probability with reward signals. RGD treats the base generator as a reference policy and introduces a reward model as a test-time controller, injecting reward into each decoding step to reshape the search trajectory without retraining the generator. Extensive offline and online experiments demonstrate the effectiveness of our approach for aligning personalization and business value. RGD has been deployed on the Kuaishou platform, bringing consistent improvements in real-world recommendation scenarios.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑