与什么相比?面向大语言模型生成基础设施即代码的人类锚定安全基准
Compared to What? A Human-Anchored Security Benchmark for LLM-Generated Infrastructure-as-Code
浏览论文内容
中文总结 AI 辅助
该研究推出GenIaC-SecBench基准,对比大语言模型与人类生成IaC的漏洞密度,发现模型漏洞密度为人类的3.21-3.87倍,供应商扩展思维表现优于提示思维链,还得出可部署性与漏洞无关联等结论,相关代码数据已公开。
中文摘要 AI 辅助
大语言模型越来越多地被用于生成基础设施即代码(IaC),单个不安全的默认设置可能会直接部署到生产环境中。此前的评估仅报告模型生成IaC的原始漏洞数量,但由于缺少人类基准,无法确定模型是否真的比工程师差。我们推出GenIaC-SecBench,这是一个按架构复杂性分层的100个部署场景的基准,对来自四个供应商的12种模型配置进行评估,生成1196个IaC制品并由三个独立策略引擎(Checkov、Trivy、KICS)扫描。关键的是,我们还使用相同的工具链扫描了634个人工编写的IaC模板,提供了第一个规模匹配的人类安全基准。漏洞密度与制品规模呈强负相关(斯皮尔曼ρ=-0.55,p<10^-77),这意味着不匹配的比较衡量的是规模而非安全性。在声明资源数量匹配的情况下,所有模型配置的漏洞密度均处于人类的3.21倍至3.87倍之间,且差距在更简单的任务中会扩大(1个资源时为4.9倍,20个或更多资源时为1.4倍)。我们将推理分解为标准生成、提示工程思维链和供应商扩展思维API。供应商扩展思维的表现明显优于提示思维链(-12.0%,p=0.0013),而提示思维链与标准生成无显著差异(-1.3%,无统计学意义)。令牌检测显示扩展思维使用的令牌不足输出预算的1%,这解释了其有限的效果。还出现了两个负面结果:可部署性与漏洞无相关性(r=0.158,p=0.625),且经典的完整案例Friedman检验不适用于实际的基准设计,这推动了Skillings-Mack统计量的应用。所有代码、数据和再生脚本均已公开。
英文摘要
Large language models increasingly author Infrastructure-as-Code (IaC), where one insecure default is provisioned straight into production. Prior evaluations report vulnerability counts for models only, and so cannot say whether models are worse than the engineers they assist. We present GenIaC-SecBench: 100 deployment scenarios across 12 model configurations from six vendors, open and closed weights, yielding 1,196 artifacts scanned by three policy engines (Checkov, Trivy, KICS) at complete coverage. Crucially we scan 634 human-authored IaC templates with the identical toolchain, giving the first size-matched human security baseline for this task. Vulnerability density is strongly inverse to artifact size (Spearman $ρ=-0.55$, $p<10^{-77}$), so unmatched comparisons measure size, not security. Size-matched, every configuration exceeds the human baseline at $3.21\times$ to $3.87\times$, and the gap widens as tasks get simpler ($4.9\times$ at one resource, $1.4\times$ at twenty or more). A majority of scenarios prescribe a security state rather than specifying function alone, so we stratify by prompt class: pooled the gap is $3.50\times$, and excluding every scenario that explicitly requests an insecure configuration still leaves all configurations above baseline ($2.4\times$ to $4.2\times$). The corpus cannot isolate unprompted default posture, and we say so. Decomposing "reasoning" into standard generation, prompted chain-of-thought, and vendor extended-thinking APIs, extended thinking beats prompted CoT ($-12.0\%$, $p=0.0013$) while prompted CoT alone is indistinguishable from standard ($-1.3\%$, n.s.); it consumes under $1\%$ of the output budget, bounding the effect. Two negative results: more deployable models are not more vulnerable ($r=0.158$, $p=0.625$), and complete-case Friedman is uncomputable here, motivating Skillings-Mack. All code and data are released.