TerraRepair:一种用于基础设施即代码修复的工具接地大语言模型代理
TerraRepair: A Tool-Grounded LLM Agent for Infrastructure-as-Code Repair
AI总结:
研究探讨工具接地能否改进基于LLM的Terraform修复及何时升级问题。提出TerraRepair工具接地LLM代理,通过检索上下文等操作进行修复。实验表明其能显著提高扫描器验证修复率,凸显工具接地改进修复效果,同时指出特定部署上下文是自主修复的主要限制。
AI中文摘要:
背景:基础设施即代码(IaC)扫描器在部署前检测Terraform和其他IaC语言中的云配置错误,但修复标记的配置大多仍需手动操作。基于大语言模型(LLM)的修复方法能修复部分问题,但可能产生不存在的结构或未解决问题。目的:研究工具接地能否改进基于LLM的Terraform修复,以及在缺少特定部署上下文时何时应升级问题。方法:提出TerraRepair,一种用于Terraform修复的工具接地LLM代理原型,具有结构化升级功能。它从Terraform引用中检索依赖上下文,参考已安装的提供程序模式,并在返回候选修复前重新运行扫描器。缺少所需上下文时,TerraRepair会升级而非编造似是而非的修复。结果:使用Checkov和Trivy这两种IaC安全扫描器,在两个故意存在漏洞的Terraform存储库上评估该工具,涵盖AWS、Azure和GCP。在AWS综合基准测试中,与受控的一次性基线相比,TerraRepair将Checkov的扫描器验证修复率从26.6%提高到78.4%,将Trivy的修复率从44.8%提高到72.4%。在多数投票协议下,其修复被标记为正确。结论:这些新结果表明,工具接地可显著改进基于扫描器验证的、基于LLM的IaC修复,但缺少特定部署上下文仍是完全自主的主要知识边界。
英文摘要:
Background: Infrastructure-as-Code (IaC) scanners detect cloud misconfigurations in Terraform and other IaC languages before deployment, but repairing the flagged configurations remains largely manual. Recent Large Language Model (LLM)-based repair approaches can repair some findings, but may hallucinate unsupported constructs or suppress warnings without fixing the issue. Aims: We study whether tool grounding can improve LLM-based Terraform repair, and when a finding should be escalated because the required deploymnet-specific context is not availble. Method: We present TerraRepair, a prototype of a tool-grounded LLM agent for Terraform repair with structured escalation. TerraRepair retrieves dependency context from Terraform references, consults the installed provider schema, and re-runs the scanner before returning a candidate repair. Then teh required context is absent, TerraRepair escalates instead of fabricating a plausible fix. Results: We evaluate our tool on two vulnerable-by-design Terraform repositories using two IaC security scanners, Checkov and Trivy, across AWS, Azure, and GCP. On the combined AWS benchmark, TerraRepair improves scanner-verified fix rates from 26.6% to 78.4% on Checkov and from 44.8% to 72.4% on Trivy, compared with a controlled one-shot baseline. It repairs are labelled as correct under a majority-vote protocol. Conclusions: These emerging results show that tool grounding can substantially improve scanner-verified LLM-based IaC repair on the studied benchmarks, while missing deployment-specific context remains the main knowledge boundary for full autonomy.