arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.11669cs.LGcs.AIcs.CL

规则丢弃(Rubric Dropout):缓解以规则为奖励的强化学习中奖励黑客行为的简单方法

Rubric Dropout: A Simple Way to Mitigate Reward Hacking in Rubric-as-Reward RL

Minglai Yang, Xinyu Guo, Utkarsh Tyagi, Mian Zhang, Razvan Dumitru, Sunjie Hou, Yunzhong He, Daniel Yue Zhang, Ying Liu

首次发表
浏览论文内容

中文总结 AI 辅助

该研究针对以规则为奖励的强化学习中的奖励黑客行为,提出Rubric Dropout方法,在医学和科学规则上训练Qwen3-8B,可提升分布外基准的黄金评判分数、降低黑客指标,30-50%丢弃比例效果最优。

中文摘要 AI 辅助

针对规则(由大语言模型评判器评分的标准列表)的强化学习已成为对无确定性答案的任务进行后训练语言模型的标准方法。然而,规则只是质量的固定代理,绝非质量的完整描述,长期针对规则训练的策略会学会利用这种差异。我们直接对此进行了测量:使用组相对策略优化(GRPO)在医学和科学规则上训练Qwen3-8B,并使用训练评判器和更强的黄金评判器对分布外(OOD)基准进行评分,发现训练过程中两个分数出现分化——训练评判器的分数持续上升,而黄金评判器的分数先达到峰值后下降,HealthBench-Hard数据集上下降3分,ResearchQA数据集上下降22分。具有固定偏差的评判器只会使黄金曲线产生恒定偏移,不会在训练分数上升时使其下降,因此这种分化是奖励黑客行为,而非评判器噪声。我们提出Rubric Dropout,一种源自神经元丢弃的一行修复方法:每一步计算奖励前,随机丢弃规则的一部分标准,使策略从未两次优化相同的规则;丢弃的子集在每个rollout组间共享,因此GRPO的组相对优势保持可比,且评估始终使用完整规则。在两个基准对中,将无丢弃、丢弃30%和50%的情况对比,丢弃在每个匹配检查点均提高了OOD黄金分数(HealthBench-Hard提高1-2分,ResearchQA提高6-7分),降低了我们跟踪的两种黑客指标,且在域内无损失;对丢弃比例的扫描显示,30-50%是广泛的最优区间,而自然替代方案(按标准对训练的有用性重新加权)在我们的设置中表现比完全不干预更差。

英文摘要

Reinforcement learning against rubrics, lists of criteria graded by an LLM judge, has become a standard way to post-train language models on tasks with no deterministic answer. The rubric, however, is a fixed proxy for quality, never a complete description of it, and a policy trained against it long enough will learn to exploit the difference. We measure this directly. Training Qwen3-8B with Group Relative Policy Optimization (GRPO) on medical and science rubrics and grading out-of-distribution (OOD) benchmarks with both the training judge and a stronger gold judge, we find that the two scores diverge during training. The training judge's score keeps climbing while the gold judge's score peaks and then falls, by 3 points on HealthBench-Hard and by 22 points on ResearchQA. A judge with a fixed bias would shift the gold curve by a constant, not send it down while the training score rises, so the divergence is reward hacking, not judge noise. We propose Rubric Dropout, a one-line fix borrowed from neuron dropout. At every step, we randomly drop a subset of the rubric's criteria before computing the reward, so the policy never optimizes the same rubric twice. The dropped subset is shared across each rollout group, so GRPO's group-relative advantages stay comparable, and evaluation always uses the full rubric. Comparing no dropout against dropout at 30% and 50% on both benchmark pairs, dropout raises the OOD gold score at every matched checkpoint (+1 to +2 points on HealthBench-Hard, +6 to +7 points on ResearchQA), lowers the two hacking measures we track, and costs nothing in domain. Sweeping the dropout fraction shows a broad 30-50% sweet spot, while the natural alternative, reweighting criteria by how useful they are to training, performs worse than no intervention at all in our setting.

发表机构

  • Scale AI
  • University of Arizona(亚利桑那大学)
  • University of Texas at Dallas(德克萨斯大学达拉斯分校)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑