arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.17776cs.LG

辩论训练可减少RLAIF中的奖励黑客行为

Debate Training Reduces Reward Hacking in RLAIF

Zachary Kenton, Lili Janzer, Rory Greig, Tian Huey Teh, Kirill Tyshchuk, Jonah Brown-Cohen, Harri Edwards, Senthooran Rajamanoharan, Noah Y. Siegel, Natasha Jaques, Rohin Shah

首次发表
浏览论文内容

中文总结 AI 辅助

该研究提出用辩论训练替代RLAIF基线,在数学任务中可减少奖励黑客行为,维持裁判性能并恢复45%的性能差距,还验证了多方面相关实验结论。

中文摘要 AI 辅助

我们证明,使用辩论(一种由较弱大语言模型(LLM)裁判裁决的生成器与评判者之间的双人对抗游戏)对大语言模型(LLM)进行RL微调,与AI反馈强化学习(RLAIF)基线相比,可减少奖励黑客行为。奖励黑客行为是RLAIF中的核心障碍:随着训练推进,策略会学习利用其AI裁判的系统性错误,导致任务性能下降,而当裁判弱于策略时,该问题会加剧,这一设置与监督日益强大的AI系统高度相关。我们研究数学任务,其最终答案正确性可验证,这使我们能够测量奖励黑客行为动态。我们训练了一个Gemini 2.5 Flash级策略,搭配冻结的、较弱的Gemini 2.5 Flash Lite裁判,将单人RLAIF基线与辩论进行对比。基线会迅速对裁判进行黑客攻击,而辩论在整个训练过程中保持裁判性能,从而在多个RL步骤中维持更高的峰值验证准确率(恢复了45%的性能差距)。额外实验表明:1)进一步削弱裁判会加快黑客攻击速度,但可通过增加一轮辩论来弥补;2)辩论激励会覆盖提示的对齐偏差;3)使用LLM裁判的RL与可验证奖励的RL相比,训练/验证奖励差距更小;4)学习使用真实标签进行批判以说服裁判是可能的,但速度较慢。总体而言,我们的结果为辩论的可行性提供了积极更新,同时强调平衡多智能体训练至关重要:若没有玩家约束,对抗训练可能会默认出现评判者的裁判黑客行为。我们表明,批判词限制(有效至150词)可成功平衡游戏并避免裁判黑客行为,尽管这会因限制评判者的表达清晰度而引入权衡。

英文摘要

We demonstrate that RL finetuning an LLM using debate, a two-player adversarial game between a generator and a critic adjudicated by a weaker LLM judge, reduces reward hacking compared to a reinforcement learning from AI feedback (RLAIF) baseline. Reward hacking is a central obstacle in RLAIF: as training progresses, the policy learns to exploit systematic errors in its AI judge, degrading task performance, a problem that worsens precisely when the judge is weaker than the policy, the setting most relevant to overseeing increasingly capable AI systems. We study mathematics tasks, where final-answer correctness is verifiable, allowing us to measure reward hacking dynamics. We train a Gemini~2.5 Flash-class policy with a frozen, weaker Gemini~2.5 Flash Lite judge, comparing a single-player RLAIF baseline against debate. While the baseline quickly hacks the judge, debate maintains judge performance throughout training, leading to a higher peak validation accuracy (45\% performance gap recovered) that persists through many RL steps. Additional experiments show that: 1) further weakening the judge leads to faster hacking, but this can be compensated by adding an additional debate round; 2) debate incentives override prompted misalignment; 3) RL using an LLM judge has a smaller train/validation reward gap than RL from verifiable rewards; 4) learning to critique to convince the judge using ground truth labels is possible but slow. Taken together, our results are a positive update on the feasibility of debate, while highlighting that balancing multi-agent training is critical: without player constraints, adversarial training risks defaulting to critic judge-hacking. We show that critique word limits (effective up to 150 words) successfully balance the game and avoid judge hacking, though this introduces a trade-off by restricting critic expressive clarity.

↑