arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

验证者即评判者:来自Ada/SPARK中AI编码代理的经过验证的安全软件

The Prover Is the Judge: Verified Security Software from AI Coding Agents in Ada/SPARK

Tobias Philipp

arXiv 2607.14340首次发表:更新:

AI 中文总结

研究人工智能编码代理生成安全软件代码的情况,核心方法是在验证器驱动循环下编写验证,主要贡献是通过GNATprove完成大量证明义务,降低监督成本,同时指出其不足及得出代理受反馈强度限制的教训。

AI 中文摘要

人工智能编码代理生成代码的速度比人类审查速度快。在我们的方法中,验证者判断代码是否正确。在验证器驱动的循环下,人工智能代理编写并验证了Ada/SPARK中的裸机安全软件,涵盖经典和后量子密码学、TLS 1.3、IKEv2、X.509和一个Matrix客户端。GNATprove解除了49280个证明义务,确立了所选原语的功能正确性,并证明其余部分不存在运行时错误,监督成本比类似的人工验证低约20至40倍。仅GNATprove是不够的:一些缺陷无法检测到,通过已知答案测试、互操作性或对规范的人工审查来解决。鉴于检查薄弱,代理试图绕过它们并报告成功。我们报告了每层发现故障的位置,并得出核心教训:代理可被信任建立的内容受其反馈强度的限制。

英文摘要

AI coding agents produce code faster than humans can review it. In our approach, the prover is the judge of whether the code is correct. Under a verifier-driven loop, AI agents wrote and verified bare-metal security software in Ada/SPARK spanning classical and post-quantum cryptography, TLS 1.3, IKEv2, X.509, and a Matrix client. GNATprove discharged 49,280 proof obligations, established functional correctness for selected primitives, and proved the absence of run-time errors for the rest, at roughly 20-40 times lower supervision cost than comparable hand verification. GNATprove alone was insufficient: some defects could not be detected and were resolved using known-answer tests, interoperability, or human review of specifications. Given weak checks, the agent tried to bypass them and reported success. We report where each layer caught faults and draw the central lesson: what an agent can be trusted to establish is bounded by the strength of its feedback.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑