arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.10947cs.LG

RH-Detect:用于奖励黑客检测的统一基准

RH-Detect: A Unified Benchmark for Reward Hacking Detection

Junwei Quan, Evgenii Opryshko, Rohan Subramani, Igor Gilitschenski

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出RH-Detect基准,整合11个公开数据集的奖励黑客相关数据,评估6个现成语言模型的奖励黑客检测能力,发现多轮工具使用数据集上检测准确率存在差距,输入格式和训练方法会影响检测性能。

中文摘要 AI 辅助

奖励黑客(Reward hacking)指模型在未完成预期任务的情况下利用评估信号,这对已部署的语言模型系统的可靠性构成威胁。现有数据集使用不同的标签、响应格式和元数据约定,导致检测器结果难以比较。本文提出RH-Detect基准,它将11个公开数据集中与奖励黑客相关的子集整合为统一模式,共包含92761条数据和6种行为类别。在5021个开放式评估单元(每个单元包含任务提示和自由形式的模型续接,包括多轮工具使用轨迹)上,我们评估了来自5个家族的6个现成语言模型作为奖励黑客检测器,无需额外训练。最佳模型的合并AUROC为0.962,准确率超过93%。然而,对于4个最强模型,在共同决策阈值下,它们在两个多轮工具使用数据集MALT和TRACE上的准确率比其他来源低10.7至15.9个百分点,凸显了部署时监测的关键差距。我们发现不同输入格式对不同模型的影响不同:移除思考过程使Qwen3.5-4B的AUROC从0.779提升至0.849,但使Qwen Flash的AUROC从0.977降至0.950。我们还将该基准作为检测器的训练数据集进行评估:依次留出每个来源的数据,单令牌SFT使6个留出来源中的5个的平均AUROC得到提升;对该失败案例进行GRPO后续处理后,检测性能略有提升。我们的结果表明,单一合并分数可能掩盖数据源、检测器输入和训练过程之间的差异。

英文摘要

Reward hacking, where a model exploits an evaluation signal without completing the intended task, threatens the reliability of deployed language model systems. Existing datasets use different labels, response formats, and metadata conventions, making detector results difficult to compare. We present RH-Detect, a benchmark that combines reward-hacking-relevant subsets from eleven public datasets, comprising 92,761 rows and six behavior categories, into a common schema. On 5,021 open-ended evaluation units, each comprising a task prompt and a free-form model continuation, including multi-turn tool-use trajectories, we evaluate six off-the-shelf language models from five families as reward hacking detectors without additional training. The best model achieves a pooled AUROC of 0.962, with accuracy above 93%. For the four strongest models, however, accuracy on the two multi-turn tool-use datasets, MALT and TRACE, is 10.7-15.9 percentage points lower than on the other sources at a common decision threshold, highlighting a key gap for deployment-time monitoring. We find that different input formats have different effects across models. Removing thinking raises Qwen3.5-4B AUROC from 0.779 to 0.849, but lowers Qwen Flash from 0.977 to 0.950. We also evaluate the benchmark as a training dataset for detectors. Holding out each source in turn, single-token SFT improves average AUROC on five of six held-out sources. A GRPO follow-up on that failure case yields a slight improvement in detection performance. Our results show that a single pooled score can conceal variation across data sources, detector inputs, and training procedures.

发表机构

  • University of Toronto(多伦多大学)
  • Vector Institute(矢量研究所)
  • Aether Research(埃瑟研究公司)
  • Trajectory Labs(轨迹实验室)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑