arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2502.19328cs.CLcs.AI

智能体奖励建模:整合人类偏好与可验证正确性信号以实现可靠的奖励系统

Agentic Reward Modeling: Integrating Human Preferences with Verifiable Correctness Signals for Reliable Reward Systems

  • Tsinghua University(清华大学)

机构由 AI 辅助整理,请以论文原文为准。

Hao Peng, Yunjia Qi, Xiaozhi Wang, Zijun Yao, Bin Xu, Lei Hou, Juanzi Li

更新

AI总结:

提出智能体奖励建模,结合奖励模型与可验证正确性信号(事实性和指令遵循)构建RewardAgent,在基准和下游任务上显著优于传统奖励模型,并用于DPO训练提升NLP性能。

AI中文摘要:

奖励模型(RMs)对于大型语言模型(LLMs)的训练和推理时扩展至关重要。然而,现有的奖励模型主要关注人类偏好,忽视了可验证的正确性信号,而这些信号在训练LLMs方面已显示出强大的潜力。在本文中,我们提出了智能体奖励建模,这是一种奖励系统,将奖励模型与来自不同方面的可验证正确性信号相结合,以提供可靠的奖励。我们实证实现了一个名为RewardAgent的奖励智能体,它将人类偏好奖励与两个可验证信号(事实性和指令遵循)相结合,以提供更可靠的奖励。我们在现有奖励模型基准上进行了全面实验,并在真实世界的下游任务上进行了推理时的最佳n次搜索。RewardAgent显著优于普通奖励模型,证明了其有效性。我们进一步使用RewardAgent构建训练偏好对,并通过DPO目标训练一个LLM,在各种NLP基准上取得了优于传统奖励模型的性能。我们的代码已公开发布,以促进进一步研究(https://github.com/THU-KEG/Agentic-Reward-Modeling)。

英文摘要:

Reward models (RMs) are crucial for the training and inference-time scaling up of large language models (LLMs). However, existing reward models primarily focus on human preferences, neglecting verifiable correctness signals which have shown strong potential in training LLMs. In this paper, we propose agentic reward modeling, a reward system that combines reward models with verifiable correctness signals from different aspects to provide reliable rewards. We empirically implement a reward agent, named RewardAgent, that combines human preference rewards with two verifiable signals: factuality and instruction following, to provide more reliable rewards. We conduct comprehensive experiments on existing reward model benchmarks and inference time best-of-n searches on real-world downstream tasks. RewardAgent significantly outperforms vanilla reward models, demonstrating its effectiveness. We further construct training preference pairs using RewardAgent and train an LLM with the DPO objective, achieving superior performance on various NLP benchmarks compared to conventional reward models. Our codes are publicly released to facilitate further research (https://github.com/THU-KEG/Agentic-Reward-Modeling).

补充信息

↑