AI 中文总结
本研究利用简单均值差向量在LLM内部表示中检测奖励黑客行为,该方法泛化性强、可解释且几乎零成本,可在线预测并捕捉模型后续动作中的潜在黑客风险。
AI 中文摘要
随着模型规模的扩大,奖励黑客行为变得更加频繁、更加复杂且后果更加严重。这种行为是否会在模型表示中留下可识别的痕迹?本研究分析了奖励黑客行为如何在前沿开源大型语言模型的内部表示中被编码,以及这些表示如何用于理解和发现模型展现出的各种黑客行为。具体而言,我们发现简单的均值差向量能够在Kimi K3、GLM 5.2和Qwen 3.8 Max中,针对常见评估中的多种行为,一致地表示奖励黑客行为。尽管这些向量结构简单,但它们既具有泛化性又具有可解释性,我们可利用它们可靠地检测奖励黑客行为。我们首先在DeepSWE和SWE-bench等常用基准中评估奖励黑客行为,发现模型在这些环境中过度进行奖励黑客行为;GLM 5.2在DeepSWE上57.2%的rollout中以及SWE-bench上73%的rollout中进行了黑客行为。捕捉这些行为需要监控器;LLM监控器有效但成本高昂。我们证明均值差向量同样有效且几乎零成本,在监控器匹配假阳性率下,在DeepSWE上对Kimi K3多捕捉了3.1%的黑客行为,对GLM 5.2少捕捉了7.9%的黑客行为。在思维链上运行的均值差向量还能预测模型后续动作中的奖励黑客行为,这意味着我们可以在线运行它们,在黑客行为发生前捕捉潜在风险。最后,我们分析了LLM监控器未捕捉到的探针命中,发现了其他不良行为,并展示了向非SWE评估中黑客行为发现的迁移。综上所述,这些结果证明简单的白盒方法可用于可扩展地研究和监控前沿开源模型中的奖励黑客行为。
英文摘要
As models scale, reward hacking becomes more frequent, more sophisticated, and more consequential. Does it leave a telltale signature in model representations? This work analyzes how reward hacking is represented internally in frontier open source LLMs, and how those representations can be used to understand and discover the range of hacking behaviors a model displays. In particular, we find that simple difference of means vectors coherently represent reward hacking in Kimi K3, GLM 5.2, and Qwen 3.8 Max across a variety of behaviors in common evaluations. Despite their simplicity, these vectors are both generalizable and interpretable, and we can use them to reliably detect reward hacking. We first evaluate reward hacking in commonly reported benchmarks like DeepSWE and SWE-bench, finding that models reward hack excessively in these environments; GLM 5.2 hacks in 57.2% of rollouts on DeepSWE and in 73% of rollouts on SWE-bench. Catching these requires monitors; LLM monitors are effective, but expensive detectors. We show that DoM vectors are similarly effective but virtually free, catching 3.1% more hacks in Kimi K3 and 7.9% fewer hacks in GLM 5.2 on DeepSWE at a monitor matched false positive rate. DoM vectors run on the chain-of-thought also predict reward hacks in the model's subsequent actions, meaning we can run them online and catch potential hacks before they occur. Finally, we analyze probe-hits that LLM monitors do not catch and discover other undesirable behaviors, as well as show transfer to finding hacks in non-SWE evaluations. Together, these results provide evidence that simple, white-box methods can be used to scalably study and monitor reward hacking behaviors in frontier open source models