arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.09159cs.AIcs.LOcs.SE

SpecGuard:在智能体作弊之前证明任务已损坏

SpecGuard: Proving a Task Is Broken Before the Agent Cheats

Param Biyani, Krishnamurthy Dvijotham

首次发表
浏览论文内容

中文总结 AI 辅助

SpecGuard通过自动形式化任务意图与测试为Lean 4规范,用内核验证冲突并生成证书,在SWE-bench上检测72.8%冲突,提供执行前安全认证。

中文摘要 AI 辅助

随着自主编码智能体的部署日益增多,任务中意外或对抗性注入的错误规格导致危险智能体行为的风险亟需解决。先前的研究表明,面对此类任务的智能体很少标记冲突,而是选择作弊,如修改测试或硬编码预期输出,且作弊行为可能造成真实损害,例如删除安全防御以使损坏的测试通过。目前尚不清楚此类冲突是否能在智能体行动之前通过独立可验证的证据确立。我们提出SpecGuard,用于检测并正式认证任务意图与测试之间的冲突。仅给定任务描述和代码库,SpecGuard将预期行为自动形式化为Lean 4规范。测试被独立形式化,Lean内核检查是否存在任何实现能同时满足两个形式化,若无则生成机器可检查的证书。在冲突的SWE-bench任务上,SpecGuard检测到高达72.8%的冲突,并正式认证高达51.1%,其冲突漏检率比基于模型的判断低近五倍。SpecGuard提供执行前安全检查,通过正式认证任务级冲突来识别奖励黑客机会,无需观察任何智能体行为。我们的代码可在该https URL获取。

英文摘要

As autonomous coding agents get increasingly deployed, the risk that accidental or adversarially injected misspecifications in tasks lead to dangerous agent behavior is critical to address. Prior work has shown that agents given such tasks rarely flag the conflict and instead cheat, editing tests or hard-coding expected outputs, and the actions taken to cheat can cause real damage, such as deleting a security defense to make a corrupted test pass. It remains unclear whether such conflicts can be established with independently verifiable evidence before the agent acts. We present SpecGuard, which detects and formally certifies these conflicts between task intent and tests. Given only the task description and codebase, SpecGuard autoformalizes the intended behaviour into a Lean 4 specification. The tests are formalized independently, and the Lean kernel checks whether any implementation could satisfy both formalizations, producing a machine-checked certificate when none can. On conflicted SWE-bench tasks, SpecGuard detects up to 72.8% of conflicts and formally certifies up to 51.1%, with a nearly five-fold lower conflict miss rate than model-based judgment. SpecGuard provides a pre-execution safety check that identifies reward-hacking opportunities through formal certification of task-level conflicts, before any agent behavior is observed. Our code is available at https://github.com/prmbiy/specguard.

发表机构

  • Google DeepMind(谷歌DeepMind)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑