arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

得分更高,回答更差:通过协议级评分标准缓解基于评分标准的强化学习中的奖励黑客行为

Scoring Higher, Answering Worse: Mitigating Reward Hacking in Rubric-Based RL via Protocol-Level Rubrics

Maoqi Liu, Junwei He, Bowen Zhang, Feiran Li, Wentao Ma, Rongyi Lin, Shuhan Zhong, Quan Fang

arXiv 2609.38847首次发表:更新:

发表机构

Beijing University of Posts and Telecommunications; ByteDance(北京邮电大学; 字节跳动)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对基于评分标准的强化学习中加法聚合导致奖励黑客的问题,提出协议级评分标准(ProRubric),通过分组聚合提升临床咨询的适当性10.8点且不损失覆盖率。

AI 中文摘要

基于评分标准的强化学习(Rubric-RL)在没有验证器的情况下训练语言模型。一个评判者检查评分标准的每个标准,并将判定结果聚合成奖励,通常通过加权求和。我们表明这种加法聚合是薄弱环节。在求和下,标准相互补偿:一个策略错过了关键决策,却可以通过提供无人要求的建议来买回分数。在临床咨询中,这样的策略得分更高但回答更差。评分标准覆盖率上升,而在保留的医生标准上的适当性却低于未训练的模型。医学标准不应受到指责。将这些标准分组,使它们必须共同成立,相同的标准,一字不变,恢复了三分之一的损失;更短的答案几乎没有恢复。因此,我们提出了协议级评分标准(ProRubric),它保留了标准所要求的内容,并改变了它们的聚合方式。它将检查表分组为几个协议级维度。一个维度仅当其所有标准都成立且其失败条款未触发时才计数。分组一次性离线完成,优化器保持不变。ProRubric将适当性提高了10.8个百分点,且不损失覆盖率,并在两种规模下均具有最佳的七个基准平均值。奖励的有效性不仅取决于评分标准验证了什么,还取决于它如何聚合。代码可在以下网址获取:此 https URL

英文摘要

Rubric-based reinforcement learning (Rubric-RL) trains language models where no verifier exists. A judge checks each criterion of a rubric, and the verdicts are aggregated into a reward, most often by a weighted sum. We show that this additive aggregation is the weak point. Under a sum, criteria compensate for one another: a policy that misses the one decision that matters can buy the points back with advice nobody asked for. On clinical consultation, such a policy scores higher and answers worse. Rubric coverage rises while appropriateness on held-out physician criteria falls below the untrained model. The medical criteria are not to blame. Grouped so that they must hold together, the same criteria, unchanged to the word, recover a third of the loss; shorter answers recover almost none. We therefore propose Protocol-level Rubrics (ProRubric), which keeps what the criteria ask for and changes how they are aggregated. It groups a checklist into a few protocol-level dimensions. A dimension counts only when all of its criteria hold and its failure clause does not fire. The grouping is done once, offline, and leaves the optimizer unchanged. ProRubric raises appropriateness by 10.8 points without losing coverage and has the best seven-benchmark average at both scales. Reward validity is set not only by what a rubric verifies, but by how it aggregates. Code is available at https://github.com/Estrellajer/ProRubric

CommentsUnder Review

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑