arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.33221cs.LG

RMB:奖励模型提升缓解奖励黑客问题

RMB: Reward Model Boosting Mitigates Reward Hacking

Jiabin Fan, Dezhi Ye, Yongchang Hao, Lili Mou

AI总结:

提出奖励模型提升(RMB)方法,通过训练多样化奖励模型并学习轻量级聚合器,增强RLHF中奖励信号的鲁棒性,有效缓解奖励黑客问题并提升对齐性能。

AI中文摘要:

从人类反馈中进行的强化学习(RLHF)是一种将大型语言模型(LLMs)与人类偏好对齐的强大技术。然而,由于代理奖励模型的不完美性,它常常遭受奖励黑客问题,即策略优化提升了代理奖励模型的得分,但实际上相对于真实人类偏好而言性能却有所下降。为了解决这一问题,我们提出了奖励模型提升(RMB),一种新颖的方法,旨在增强RLHF中奖励信号的鲁棒性和可靠性。RMB首先训练一组带有多样性促进正则化器的奖励模型,这鼓励每个模型学习奖励景观的互补方面。然后,RMB依据提升原则学习一个轻量级聚合器,将多样化奖励模型的输出聚合为更准确且更稳健的奖励信号。我们的大量实验表明,RMB在分布内和分布外数据集上均显著提高了奖励准确性,大幅缓解了奖励黑客问题,并最终提升了RLHF的性能。

英文摘要:

Reinforcement Learning from Human Feedback (RLHF) is a powerful technique for aligning large language models (LLMs) with human preference. However, it often suffers from the reward hacking issue, where policy optimization improves the proxy reward model while actually degrading performance with respect to the true human preference, due to the imperfection of the proxy. To address this, we propose Reward Model Boosting (RMB), a novel approach that enhances the robustness and reliability of the reward signal for RLHF. RMB first trains a set of reward models with a diversity-promoting regularizer. This encourages each model to learn complementary aspects of the reward landscape. Then, RMB learns a lightweight aggregator in the principle of boosting to aggregate the outputs of the diverse reward models into a more accurate and robust reward signal. Our extensive experiments demonstrate that RMB significantly improves reward accuracy on both in-distribution and out-of-distribution datasets, substantially mitigating the reward hacking issue and ultimately improving RLHF performance.

补充信息

↑