可验证黑客终端基准:评估终端任务中的奖励黑客行为
Hack-Verifiable Terminal Bench: Evaluating Reward Hacking in Terminal Tasks
浏览论文内容
中文总结 AI 辅助
本研究提出可验证黑客终端基准(HVTB),将可验证黑客环境(HVE)适配至终端基准,用于测量前沿模型的奖励黑客率,并探究提示能否缓解该行为。
中文摘要 AI 辅助
随着智能体变得更加强大且自主,其奖励黑客(满足任务的检查要求却违背任务意图)的倾向成为日益重要的失败模式。测量奖励黑客本身颇具挑战性,因为检测通常依赖人工检查或大语言模型(LLM)评判,两者均可能不可靠。可验证黑客环境(HVE)方法通过在任务中嵌入可检测的黑客行为,实现自动且可靠的奖励黑客识别。本研究将HVE适配至终端基准(Terminal Bench,一个领先的真实世界终端与编码任务基准),并提出可验证黑客终端基准(HVTB)。利用HVTB,我们测量前沿模型的奖励黑客率,研究包含不同程度黑客信息的提示能否缓解该行为,进而测试提示能否不仅预防已知的奖励黑客策略,还能应对提示未预见的“未知未知”漏洞。我们在指定网址发布所有环境与智能体轨迹。
英文摘要
As agents grow more capable and autonomous, their tendency to reward hack, satisfying a task's checks while violating its intent, becomes an increasingly important failure mode. Measuring reward hacking is itself challenging, as detection typically relies on human inspection or LLM judges, both of which can be unreliable. The hack-verifiable environments (HVE) methodology addresses this challenge by embedding detectable hacks into tasks, allowing reward hacks to be identified automatically and reliably. In this work, we adapt HVE to Terminal Bench, a leading benchmark of real-world terminal and coding tasks, and introduce Hack-Verifiable Terminal Bench (HVTB). Using HVTB, we measure reward-hacking rates across frontier models and study whether prompts with varying amounts of information on the hack can mitigate this behavior. This lets us test whether prompting can prevent not only known reward-hacking strategies, but also 'unknown unknown' exploits that the prompt does not anticipate. We release all environments and agent traces at https://majoroth.github.io/hack-verifiable-environments/hvtb
发表机构
- Tel Aviv University(特拉维夫大学)
- University of California Santa Barbara(加州大学圣巴巴拉分校)
机构由 AI 辅助整理,请以论文原文为准。