arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

在执行点应用设计安全:受管控的安全需求如何影响AI生成代码的安全性

Applying Security by Design at the Point of Execution: How Governed Security Requirements Affect the Security of AI-Generated Code

Pedro Farinha

arXiv 2610.10659首次发表:更新:

发表机构

Shiftleft - Secure Software Engineering, Lda.(Shiftleft - Secure Software Engineering, Lda.)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究在执行点通过MCP服务器交付受管控的SbD-ToE安全需求,可提升AI生成代码的安全性,在两个基准测试中均取得显著效果,但存在研究局限性,建议测量代码与适用需求的一致性。

AI 中文摘要

设计安全要求在编写代码前就定义安全需求,而安全代码基准通常衡量相反的情况:智能体在未收到需求的情况下接收任务,并使用其未见过的安全测试进行评分。我们测量了当从受管控的设计安全知识库(SbD-ToE)中选择的安全需求通过模型上下文协议(MCP)服务器在执行点提供给生成器时,会发生什么变化。在使用语言模型评判的DualGauge上,我们运行了59个Python任务:通过所有安全测试的任务比例从44.1%上升到78.0%,通过的安全测试比例从77.4%上升到93.0%。在使用容器中运行功能测试和真实漏洞的BaxBench上,我们运行了28个后端场景:在功能正确的解决方案中,无成功漏洞的比例从65%上升到86%,与作者的Oracle Security Reminder(85%,他们称之为不切实际的上限)相当;需求交付在不知道测试的情况下达到了该上限。该研究存在局限性:预先指定的联合安全与功能指标在两个基准上均未显著改善;使用受管控需求编写的代码更严格,且在修复任务从未说明的值时未通过功能测试,我们在事后分析中逐例归因这些失败;每个实验使用一个模型,每个任务生成一次;需求来自作者的开源手册,我们不评估其完整性。在执行点交付的受管控安全需求似乎提高了生成代码的安全性,但当前基准无法判断该代码是否符合适用的安全需求,仅能判断其是否显示评估者构建来检测的弱点,我们建议测量与适用需求的一致性。

英文摘要

Security by design asks that security requirements are defined before code is written. Secure-code benchmarks typically measure the opposite situation: the agent receives the task without requirements and is scored with security tests it has not seen. We measured what changes when security requirements, selected from a governed security-by-design knowledge base (SbD-ToE) and delivered through a Model Context Protocol (MCP) server, are given to the generator at the point of execution. On DualGauge, which scores with a language-model judge, we ran 59 Python tasks: the share of tasks passing all security tests rose from 44.1% to 78.0% and the share of security tests passed from 77.4% to 93.0%. On BaxBench, which runs functional tests and real exploits in containers, we ran 28 backend scenarios: among functionally correct solutions, the share with no successful exploit rose from 65% to 86%, comparable to the authors' Oracle Security Reminder (85%), which they call an unrealistic upper bound; the requirement delivery reached it without knowing the tests. The study has limits. The pre-specified joint secure-and-functional metric did not improve significantly on either benchmark. The code written with the governed requirements is stricter, and it failed functional tests that fix values the task never states; we attribute those failures case by case in a post-hoc analysis. Each experiment used one model and one generation per task. The requirements come from the author's open-source manual; we do not assess its completeness. Governed security requirements delivered at the point of execution appear to improve the security of generated code. Current benchmarks cannot tell whether that code meets the security requirements that apply to it, only whether it shows the weaknesses their evaluators were built to detect. We propose measuring conformance with the applicable requirements.

Comments17 pages, 3 figures, 4 tables, 7 appendices. Artefact: doi:10.5281/zenodo.23213449 and doi:10.6084/m9.figshare.34168182

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑