评估与防范AI生成Ansible代码中的安全隐患
Evaluating and Preventing Security Smells in AI-Generated Ansible Code
- School of Computing and Mathematical Sciences, University of Waikato(怀卡托大学计算与数学科学学院)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
该研究评估AI生成Ansible代码的安全隐患,提出扩展CO-STAR框架的方法,使4/16模型生成合规代码,领先模型CIS合规率达95%-100%,代码质量提升19%-49%。
AI中文摘要:
AI编码助手生成基础设施即代码(Infrastructure as Code),但尚无研究检验此类代码是否符合安全要求。这一问题至关重要,因为基础设施代码中的安全隐患会传播至已部署系统,导致基础设施存在安全漏洞且不可信。我们评估了16款AI模型生成Apache Tomcat v10和MongoDB v7的Ansible角色的情况,针对CIS基准对278个Ansible角色进行分析。在无安全指导的情况下,16款AI模型生成的代码均存在安全隐患,导致基础设施存在漏洞,无法通过合规性验证,性能逊于人类开发者编写的代码。我们提出一种方法,通过扩展CO-STAR框架将Ansible最佳实践和CIS基准整合到提示词中,可在代码生成阶段防范安全隐患,而非部署后才检测。应用该方法后,16款模型中有4款生成符合要求的代码,其中领先模型达到95%-100%的CIS合规性,是人类开发者23%-43%合规率的四倍,整体代码质量提升19%-49%。其余12款模型失败并非因无法生成代码,而是无法遵循含多重约束的指令。对于具备能力的模型,该方法无需重新训练,可通过系统提示词实现。
英文摘要:
AI coding assistants generate Infrastructure as Code, yet no work has examined whether this code meets security requirements. This matters because security smells in infrastructure code propagate to deployed systems, producing infrastructure that is insecure and untrustworthy. We evaluate 16 AI models generating Ansible roles for Apache Tomcat v10 and MongoDB v7, analysing 278 Ansible roles against CIS benchmarks. Without security guidance, all 16 AI models produced code containing security smells, resulting in vulnerable infrastructure that fails compliance verification and underperforms code written by human developers. We introduce an approach integrating Ansible best practices and CIS benchmarks into prompts through an extended CO-STAR framework, enabling security smell prevention during synthesis rather than detection after deployment. When this approach is applied, 4 out of 16 models generate compliant code, with the leading model achieving 95%-100% CIS compliance, a fourfold improvement over humans at 23%-43%, with overall code quality improving by 19%-49%. The remaining 12 models fail not because they cannot generate code but because they cannot follow instructions with multiple constraints. For capable models, the approach requires no retraining and can be adopted through system prompts.