RRC:基于排序的奖励构造解锁LLM强化学习中的生成式奖励模型
RRC: Unlocking Generative Reward Models in LLM Reinforcement Learning via Ranking-Based Reward Construction
浏览论文内容
中文总结 AI 辅助
本研究针对生成式奖励模型在LLM强化学习中潜力未充分发挥的问题,提出RRC方法,通过两种互补策略提升RL训练效果,在多基准上取得一致增益。
中文摘要 AI 辅助
奖励建模的最新进展展现出从判别式奖励模型到生成式奖励模型的范式转变。然而,尽管生成式奖励模型在响应排序方面能力强大,但其在强化学习(RL)中的潜力尚未得到充分发挥。我们的分析表明,这一局限源于生成式奖励建模的比较性质与现有RL算法采用的标量评分范式之间存在不匹配。为弥合这一差距,我们提出了基于排序的奖励构造(Ranking-based Reward Construction,RRC)方法,该方法通过从相对偏好排序中推导奖励,使生成式奖励模型能够提供更有效的RL学习信号。RRC引入了两种互补策略:利用采样响应间比较的自竞争排序,以及通过少量参考响应实现可扩展的基于排序的奖励构造的锚点引导排序。在开放式对话和推理基准上开展的实验表明,RRC显著提升了生成式奖励模型的RL训练效果,相较于现有奖励构造方法取得了一致的性能增益。我们的代码可在该https URL获取。
英文摘要
Recent advances in reward modeling show a paradigm shift from discriminative reward models to generative reward models. However, despite their strong capabilities in response ranking, generative reward models have not realized their potential in reinforcement learning (RL). Our analysis reveals that this limitation arises from a mismatch between the comparative nature of generative reward modeling and the scalar scoring paradigm adopted by existing RL algorithms. To bridge this gap, we propose a Ranking-based Reward Construction (RRC) approach, which enables generative reward models to provide more effective RL learning signals by deriving rewards from relative preference rankings. RRC introduces two complementary strategies: self-competitive ranking, which exploits comparisons among sampled responses, and anchor-guided ranking, which enables scalable ranking-based reward construction with a small set of reference responses. Experiments across open-ended chat and reasoning benchmarks demonstrate that RRC substantially improves RL training with generative reward models, achieving consistent gains over existing reward construction approaches. Our code can be found at https://github.com/wangclnlp/RRC.
发表机构
- School of Computer Science and Engineering, Northeastern University(东北大学计算机科学与工程学院)
- NiuTrans Research(NiuTrans研究院)
- Institute of Psychology, CAS(中国科学院心理研究所)
- Kunming University of Science and Technology(昆明理工大学)
机构由 AI 辅助整理,请以论文原文为准。