arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

IaC-Guard-V:面向LLM生成的基础设施即代码修复的验证框架

IaC-Guard-V: A Verification Framework for LLM-Generated Infrastructure-as-Code Repairs

Lokesh Chauhan

arXiv 2609.28488首次发表:更新:

AI 中文总结

IaC-Guard-V提出一个四维验证框架,评估LLM生成的IaC修复,基于70个真实工件和630次运行,发现验证引导迭代修复可将修复率从32-50%提升至68-92%,且开源模型成本仅为商业模型的十二分之一。

AI 中文摘要

基础设施即代码(IaC)配置错误是云安全事件的主要原因,大型语言模型(LLMs)正日益被提出作为自动化修复代理。然而,LLM生成的IaC修复的可信度仍未得到充分探索。IaC提出了独特的验证挑战:安全扫描器在不同工具和配置下可能表现不同,基础设施语义无法通过普通单元测试验证,且特定于提供商的规则造成了碎片化的验证环境。我们提出了IaC-Guard-V,一个以验证为中心的框架,通过四个维度评估AI生成的IaC修复:语法有效性、目标问题解决、回归安全性和补丁最小性。我们构建了一个包含70个真实世界错误配置的Terraform和Kubernetes工件的基准,涵盖70条独特的扫描器规则和八类违规,并在三个LLM家族中评估了三种修复策略,共630次运行。尽管所有模型都实现了100%的语法有效性,但只有32-50%的单次修复通过完整验证。验证引导的迭代修复显著将验证修复率提高到68-92%。一个使用验证引导修复的开源模型以每验证修复成本十二分之一的价格优于未使用验证的最强商业模型。Kubernetes修复在商业模型中达到接近完美的比率,但开源模型需要验证引导迭代。结构化提示始终降低Terraform修复质量,挑战了关于受限LLM输出的常见假设。基准和工件已发布以供复现。

英文摘要

Infrastructure-as-Code (IaC) misconfigurations are a leading cause of cloud security incidents, and Large Language Models (LLMs) are increasingly proposed as automated repair agents. Yet the trustworthiness of LLM-generated IaC repairs remains underexplored. IaC presents distinctive verification challenges: security scanners may behave differently across tools and configurations, infrastructure semantics cannot be validated through ordinary unit tests, and provider-specific rules create a fragmented verification landscape. We present IaC-Guard-V, a verification-centered framework that evaluates AI-generated IaC repairs through four dimensions: syntactic validity, target-issue resolution, regression safety, and patch minimality. We construct a benchmark of 70 real-world misconfigured Terraform and Kubernetes artifacts spanning 70 unique scanner rules and eight violation classes, and evaluate three repair strategies across three LLM families in 630 runs. Although all models achieve 100% syntactic validity, only 32-50% of single-shot repairs pass full verification. Verification-guided iterative repair significantly improves verified-fix rates to 68-92%. An open-source model with verification-guided repair outperforms the strongest commercial model without verification at one-twelfth the cost per verified fix. Kubernetes repairs reach near-perfect rates for commercial models but require verification-guided iteration for the open-source model. Structured prompting consistently reduces Terraform repair quality, challenging common assumptions about constrained LLM output. The benchmark and artifacts are released for reproducibility.

Comments11 pages, 3 figures. Pre-peer-review manuscript. Accepted as a regular paper at QRS 2026. Code and replication artifacts: https://github.com/lokesh0186/iac-guard-v . This arXiv version predates the post-review and camera-ready revisions

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑