Mitigating Reward Hacking in RLHF via Bayesian Non-negative Reward Modeling
通过贝叶斯非负奖励建模缓解RLHF中的奖励黑客
机构 * Zhejiang University(浙江大学)
AI总结 提出贝叶斯非负奖励模型(BNRM),通过非负因子分析和变分推断,在Bradley-Terry偏好模型中实现解耦与去偏,有效缓解奖励过度优化,提升鲁棒性和可解释性。
Comments Accepted as an Oral presentation at ICML 2026. The code is available at https://github.com/GuoweiRong/Bayesian-Non-negative-Reward-Model