arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.10762cs.CRcs.AIcs.SE

超越静态保证:衡量安全敏感及LLM生成的Python代码中的静态通过-动态失败差距

Beyond Static Guarantees: Measuring the Static-Pass Dynamic-Fail Gap in Security-Sensitive and LLM-Generated Python Code

Jessica Pourleyli, Maitreyee Das Urmi, Glaucia Melo

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出静态通过-动态失败(SPDF)现象及三阶段智能体流水线,通过静态扫描、LLM推理和动态验证评估1,355个Python样本,发现14.53%的静态干净样本存在可利用漏洞,表明静态分析与运行时安全是分层保证而非可互换度量。

中文摘要 AI 辅助

大型语言模型(LLM)的进展推动了寻求可扩展方法来评估生成的和安全敏感的软件安全性的进程。静态分析作为一种可扩展、可复现且成本低廉的安全门控被广泛采用,但它无法直接观察运行时利用行为。依赖于对抗性输入、执行上下文或利用链的漏洞可能绕过静态检查,而在实践中仍可利用,然而通过静态分析往往被视为安全行为的证据。本文引入了静态通过-动态失败(SPDF)现象,并提出了一种三阶段智能体流水线,结合静态扫描、LLM驱动的常见弱点枚举(CWE)推理以及隔离Docker容器中的自主利用验证。我们评估了来自SecurityEval、RedCode和CyberNative数据集的1,355个Python样本。在复合Bandit-Semgrep门控下未产生任何发现的654个样本中,LLM检测阶段在235个文件中识别出394个候选漏洞。动态验证在95个文件中确认或部分确认了可利用性,得到包含性流水线率为14.53%(大约每7个静态干净样本中有1个)。该比率代表在Bandit-Semgrep干净的样本中,流水线识别出候选漏洞并获得支持可利用性的运行时证据的比例。结果因数据集而异:在候选文件-CWE对中,RedCode的确认可利用性为33.7%,CyberNative为28.6%,SecurityEval为5.4%。几个频繁确认的类别,包括CWE-338和CWE-916,既未被Bandit也未被Semgrep标记。这些发现表明,静态分析成功和运行时安全是软件保证的分层级别,而非可互换的度量,并有可能重塑AI生成和安全敏感代码的评估方式。

英文摘要

Advances in large language models (LLMs) fuel the quest for scalable methods to assess the security of generated and security-sensitive software. Static analysis is widely adopted as a scalable, reproducible, and inexpensive security gate, but cannot directly observe runtime exploit behaviour. Vulnerabilities dependent on adversarial inputs, execution context, or exploit chaining may evade static checks while remaining exploitable in practice, yet passing static analysis is often treated as evidence of secure behaviour. This paper introduces the Static-Pass Dynamic-Fail (SPDF) phenomenon and a three-stage agentic pipeline combining static scanning, LLM-driven Common Weakness Enumeration (CWE) reasoning, and autonomous exploit verification in isolated Docker containers. We evaluate 1,355 Python samples from SecurityEval, RedCode, and CyberNative datasets. Of the 654 samples producing no findings under the composite Bandit-Semgrep gate, the LLM detection stage identified 394 candidate vulnerabilities across 235 files. Dynamic verification confirmed or partially confirmed exploitability in 95 files, yielding an inclusive pipeline rate of 14.53% (roughly 1 in 7 statically clean samples). This rate represents the proportion of Bandit-Semgrep-clean samples for which the pipeline identified a candidate vulnerability and obtained runtime evidence supporting exploitability. Outcomes varied by dataset: among candidate file--CWE pairs, confirmed exploitability was 33.7% for RedCode, 28.6% for CyberNative, and 5.4% for SecurityEval. Several frequently confirmed classes, including CWE-338 and CWE-916, were flagged by neither Bandit nor Semgrep. These findings indicate that static-analysis success and runtime security are hierarchical layers of software assurance rather than interchangeable measures, and have the potential to reshape how AI-generated and security-sensitive code is evaluated.

发表机构

  • Toronto Metropolitan University(多伦多都会大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑