arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.16187cs.CRcs.AIcs.SE

保护AI生成代码:一种即时漏洞检测与修复流水线

Securing AI-Generated Code: A Just-in-Time Vulnerability Detection and Remediation Pipeline

Mikhail Surikov

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出一种自动化安全评估流水线,可检测、修复AI生成代码的漏洞,在Claude的4种模型上评估发现,P2配置表现优于P1,Sonnet 4.6的修复效果最佳,且流水线有效性与初稿安全性不同。

中文摘要 AI 辅助

AI辅助开发工具生成易受攻击代码的比例很高,但很少有自动化机制能够以开发速度检测、补充、修复和验证安全问题,尤其是那些将修复措施基于现实威胁背景的机制。本文提出一种自动化安全评估流水线,该流水线从LLMSecEval提示生成Python代码,使用CodeQL和Bandit与独立的代码验证大语言模型(Code Validator LLM)并行扫描漏洞,用MITRE ATT&CK技术、CWE观测示例和Python最佳实践指南补充代码验证模型的发现,通过代码生成大语言模型(Code Generation LLM)生成修复方案,再使用CodeQL和Bandit重新扫描以验证结果。评估了两种流水线配置:流水线1(P1)仅使用经补充的代码验证模型发现,流水线2(P2)还会接收初始的CodeQL和Bandit发现。两种配置均在Claude的4种模型(Opus 4.8、Sonnet 4.6、Sonnet 5和Haiku 4.5)上运行,针对26个LLMSecEval提示产生80次运行,覆盖9个CWE类别。P1减少了所有4种模型的静态分析器发现,范围从-9%(Opus 4.8)到-54%(Sonnet 5);P2进一步加深了这些减少,范围从-29%(Opus 4.8)到-69%(Haiku 4.5),且P2在所有模型上的表现均优于P1。所有配置的判定一致性平均约为81%的模态一致性,P2比P1略微更稳定。修复措施在15%-22%的案例中引入了新漏洞:约70%涉及单个新发现,P2减少了4种模型中3种的变更量,仅Sonnet 5是例外。值得注意的是,最佳的代码生成大语言模型(Opus 4.8)并非最佳流水线表现者,因为Sonnet 4.6在P2修复后产生的残留发现最少、通过率最高,表明流水线有效性与初稿安全性是不同的属性。

英文摘要

AI-assisted development tools generate vulnerable code at significant rates, yet few automated mechanisms exist to detect, enrich, fix, and verify security issues at development velocity, particularly ones that ground remediation in real-world threat context. This paper presents an automated security evaluation pipeline that generates Python code from LLMSecEval prompts, scans for vulnerabilities using CodeQL and Bandit in parallel with an independent Code Validator LLM, enriches the Code Validator findings with MITRE ATT&CK techniques, CWE Observed Examples, and Python best practice guidelines, generates fixes via the Code Generation LLM, and re-scans with CodeQL and Bandit to verify outcomes. Two pipeline configurations were evaluated: Pipeline 1 (P1), using enriched Code Validator findings only, and Pipeline 2 (P2), where it additionally receives the initial CodeQL and Bandit findings. Both configurations were run across four Claude models: Opus 4.8, Sonnet 4.6, Sonnet 5, and Haiku 4.5, producing 80 runs against 26 LLMSecEval prompts covering 9 CWE categories. P1 reduced static analyzer findings across all four models, ranging from -9% (Opus 4.8) to -54% (Sonnet 5). P2 deepened these reductions further, ranging from -29% (Opus 4.8) to -69% (Haiku 4.5), with P2 outperforming P1 for every model. Verdict consistency averaged approximately 81% modal agreement across all configurations, with P2 marginally more stable than P1. Remediation introduced new vulnerabilities in 15-22% of cases: roughly 70% involved a single new finding, and P2 reduced churn for three of four models, with Sonnet 5 as the sole exception. Notably, the best Code Generation LLM (Opus 4.8) was not the best pipeline performer, as Sonnet 4.6 produced the lowest residual findings and highest pass rate after P2 remediation, suggesting that pipeline effectiveness and first-draft security are distinct properties.

发表机构

  • Georgia Institute of Technology(佐治亚理工学院)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑