揭示并缓解多奖励强化学习中聚合诱导的奖励黑客行为
Uncovering and Mitigating Aggregation-Induced Reward Hacking in Multi-Reward Reinforcement Learning
浏览论文内容
中文总结 AI 辅助
针对多奖励强化学习中聚合诱导的奖励黑客问题,提出自适应多奖励投影(AMRP)方法,在三类任务中提升奖励概况平衡与下游性能,且兼容多种RL算法。
中文摘要 AI 辅助
大型语言模型的强化学习微调日益采用多个奖励维度,包括可验证规则、特定任务评估器和学习到的奖励模型,以在多样化能力上提供更丰富的监督。这些维度通常通过固定聚合权重进行标量化。我们发现一种失效模式,其中聚合本身会诱导奖励黑客行为:静态投影将性质不同的奖励概况别名为单个标量,引导优化转向奖励信号最容易、最密集或系统偏好的维度。训练过程中,这会将策略困在次优概况中,阻止其收敛到能产生更高任务性能的更均衡概况。为解决此问题,我们提出自适应多奖励投影(AMRP),这是一种轻量级在线方法,利用相对缺口、奖励波动性和近期进展三个信号重新分配聚合权重,增加滞后、不稳定或停滞维度的压力,同时减轻饱和维度的压力。在结构化推理、基于引用的生成和GRPO下的开放式对齐任务中,AMRP相比固定和动态加权基线始终提升奖励概况平衡与下游性能;其在GDPO和PPO下仍有效,支持与各类RL算法兼容。我们的代码可在this https URL获取。
英文摘要
Reinforcement learning fine-tuning of large language models increasingly adopts multiple reward dimensions, including verifiable rules, task-specific evaluators, and learned reward models, to provide richer supervision across diverse capabilities. These dimensions are commonly scalarized with fixed aggregation weights. We identify a failure mode in which aggregation itself induces reward hacking: static projection aliases qualitatively different reward profiles into a single scalar, steering optimization toward whichever dimensions are easiest, densest, or systematically favored by the reward signal. Over training, this traps the policy in suboptimal profiles and prevents convergence to better-balanced ones that would yield higher task performance. To address this, we propose Adaptive Multi-Reward Projection (AMRP), a lightweight online method that reallocates aggregation weights using three signals, relative shortfall, reward volatility, and recent progress, increasing pressure on lagging, unstable, or stagnant dimensions while relieving saturated ones. Across structured reasoning, citation-grounded generation, and open-ended alignment under GRPO, AMRP consistently improves reward-profile balance and downstream performance over fixed and dynamic weighting baselines; it also remains effective with GDPO and PPO, supporting compatibility across RL algorithms. Our code is available at https://github.com/yyhappier/AMRP.git.
发表机构
- University of Science and Technology of China(中国科学技术大学)
- Peking University(北京大学)
- Meituan(美团)
机构由 AI 辅助整理,请以论文原文为准。