安全的假象:AI与人类C++代码的多层验证
The Illusion of Safety: Multi-Tier Verification of AI vs. Human C++ Code
AI总结:
提出VULBENCH-CPP基准,通过四层验证发现AI生成C++代码的运行时违规率约为人类代码的两倍,静态分析会因代码长度产生安全假象。
AI中文摘要:
大型语言模型越来越多地生成C++代码,这是一种内存不安全的语言,一个被忽视的违规就可能成为可利用的漏洞。然而,大多数对AI生成代码的安全评估仅依赖静态分析,它只标记警告而不确认运行时违规或推理未测试的路径。我们询问AI生成的C++是否在可测量程度上比人类编写的代码更不安全,以及常见的验证工具是否对风险达成一致。我们引入了VULBENCH-CPP基准,包含来自三个开源权重LLM(Gemma 3 27B IT、LLaMA 3.3 70B Instruct、Qwen 2.5 Coder 32B Instruct)和人类作者在851个竞赛编程任务中的8,918个C++程序。每个程序通过四个验证层进行标注:功能测试、静态分析(cppcheck、clang-tidy)、动态分析(ASan/UBSan)和有界模型检查(ESBMC)。考虑到共享任务解决方案之间的相关性,我们发现AI生成的代码触发确认运行时违规的可能性大约是人类代码的两倍,即使在控制代码长度和测试通过率后也是如此。在静态分析下,两者看起来同样安全,但这具有误导性:表面相似性反映的是代码长度而非实际安全性,并且各层检测到的主要违规类别不同,因此没有单一层是足够的。这一差距在独立生成中保持一致。
英文摘要:
As large language models (LLMs) are increasingly deployed for systems programming, their ability to generate secure C++ code, where a single memory-safety failure creates an exploitable vulnerability, remains a critical concern. Yet most security evaluations of AI-generated code rely on static analysis alone, which flags warnings without confirming run- time violations or reasoning about untested paths. This study investigates whether AI-generated C++ is measurably less safe than human-written code, and whether common verification tools agree on the risk. We introduce VULBENCH-CPP, a benchmark of 8,918 C++ programs from three open-weight LLMs (Gemma 3 27B IT, LLaMA 3.3 70B Instruct, Qwen 2.5 Coder 32B Instruct) and human authors across 851 competitive-programming tasks. Each program is annotated by four verification tiers: functional testing, static analysis (cppcheck, clang-tidy), dynamic analysis (ASan/UBSan), and bounded model checking (ESBMC). Account- ing for the correlation among solutions to a shared task, we find that AI-generated code is roughly twice as likely as human code to trigger a confirmed runtime violation, even after controlling for code length and test pass-rate. Under static analysis the two look equally safe, but this is misleading: the apparent similarity reflects code length rather than real safety, and the tiers detect largely different classes of violation, demonstrating that no single tier is sufficient. These vulnerability patterns remain consistent across independent generations. We release the benchmark, harness, and annotated results.