arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.16816cs.LGcs.CL

ImpossibleRubrics:对生成式评分标准作为奖励信号的压力测试

ImpossibleRubrics: Stress-Testing Generated Rubrics as Reward Signals

发表机构新加坡国立大学 · 北京大学 · 中国科学院自动化研究所
查看机构详情
  • National University of Singapore(新加坡国立大学)
  • Peking University(北京大学)
  • Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)

机构由 AI 辅助整理,请以论文原文为准。

Bowen Qin, Yi Xie, Yesheng Liu, Xi Yang

首次发表
浏览论文内容

中文总结 AI 辅助

针对生成式评分标准作为奖励信号的鲁棒性,提出ImpossibleRubrics基准,通过对抗性测试揭示评分标准被利用源于对错误细节的具体化,而非模糊性。

中文摘要 AI 辅助

语言模型生成的评分标准越来越多地被用作基于评分标准的强化学习、LLM-as-a-judge评估和自动评分的奖励信号。此类评分标准只有在奖励诚实回答而非针对其优化利用的对抗性回答时才是可靠的。然而,它们对此类优化的鲁棒性仍然知之甚少。我们隔离了最困难的场景:不可能任务,即提示迫使模型得出无依据的结论,因此唯一诚实的回应是承认其不可能性。我们引入了ImpossibleRubrics,一个包含169个不可能任务的基准,涵盖六种不可能性类别,每个任务都配有一个可验证的预言机证书,规定诚实回答可以且不可以声称的内容,以及48个可回答的对照任务。ImpossibleRubrics不提供固定的评分标准,而是提供任务环境和证书,允许评分标准在下游生成,然后进行对抗性测试,以检验它们是否奖励违反证书的回答。在无偏的150/169环境子集上,十一个生成器被利用的概率为8%至26%;在刻意选择的压力子集上,我们测得的最强生成器仍被利用36%,而忠实于证书的评分标准被利用0%,因此我们测量的是评分标准质量差距,而非任务不可能性。一项结果与直觉相悖。一个通用的评分标准(“果断,惩罚含糊其辞”)对每个任务不加修改地使用,被利用的概率为64%,而十一个生成器中有七个在针对每个任务定制评分标准时被利用的概率更高。定制的标准似乎告诉攻击者应该捏造哪个主张。问题不在于评分标准含糊;而在于它们对错误的事情过于具体。

英文摘要

Language model-generated rubrics are increasingly used as reward signals for rubric-based reinforcement learning, LLM-as-a-judge evaluation, and automated grading. Such rubrics are reliable only if they reward honest answers over adversarial answers optimized to exploit them. Yet their robustness to such optimization remains poorly understood. We isolate the hardest regime: impossible tasks, where the prompt pressures the model toward an unsupported conclusion, so the only honest response is to acknowledge the impossibility. We introduce ImpossibleRubrics, a benchmark of 169 impossible tasks spanning six impossibility categories, each paired with a verifiable oracle certificate specifying what an honest answer may and may not claim, together with 48 answerable controls. Rather than providing fixed rubrics, ImpossibleRubrics provides task environments and certificates, allowing rubrics to be generated downstream and then adversarially tested for whether they reward certificate-violating answers. Eleven generators are exploited 8--26% of the time on the unbiased 150-of-169 environment cut; on a deliberately selected stress cut the strongest generator we measured is still exploited 36% while a certificate-faithful rubric is exploited 0%, so what we measure is a rubric-quality gap, not task impossibility. One result runs against intuition. A single generic rubric ("be decisive, penalize hedging") used unchanged for every task is exploited 64% of the time, and seven of the eleven generators are exploited more often than that while writing a rubric tailored to each one. The tailored criteria appear to tell an attacker which claim to fabricate. The problem is not that rubrics are vague; it is that they are specific about the wrong things.

↑