AI 中文总结
该研究通过分析IaC-Eval基准的5968个场景,发现迭代式LLM驱动的IaC修复会引入约3.3%场景的安全退化,主要由资源重构导致,第3次迭代是最优停止点,为安全反馈设计提供指导。
AI 中文摘要
背景:迭代反馈循环是改进大语言模型(LLM)生成的基础设施即代码(IaC)的主流范式:Checkov和terraform validate等验证器会将错误信号反馈给后续的修复尝试。现有研究报告的是累积最优指标,这类指标本质上不会下降,因此从未有人针对IaC研究过逐次迭代的原始安全性轨迹。目标:本研究探讨安全退化(即之前通过的CIS基准检查在某次修复迭代后失败),以确定迭代式LLM修复在修复其他问题时是否会、以及多频繁地降低安全性。方法:我们分析了IaC-Eval基准中的5968个场景时间线,每个场景在一种配置下最多运行5次修复迭代。15种配置(6种特定于模型的检索增强生成RAG、9种模型聚合的非RAG,每种配置有3种温度设置)产生了4440次迭代转换,其中包含两侧的Checkov数据。我们跟踪30个单独的CIS检查ID,并从代码差异中对根本原因进行分类,检测模式分为两种:标准模式(包含所有情况)和严格模式(仅排除检查失败)。结果:在标准检测模式下,13.8%的场景(24.8%的转换)至少出现一次退化;在严格检测模式下,该比例降至3.3%的场景(5.2%的转换),表明大多数明显的退化是多资源测量的人为因素。资源重构(占79.0%)是主要根本原因。退化转换显示代码变动量是正常情况的2.6倍(Cohen's d=0.90),严格模式检查波动性是正常情况的4.9倍(d=1.49)。在标准模式的退化中,36.6%会在平均1.2次迭代内自行纠正;第3次迭代是最优停止点。结论:迭代式IaC修复确实会引入安全退化,但保守且可辩护的比例约为3.3%的场景。我们的发现为安全感知型反馈循环设计和可操作的迭代预算指导提供了依据。
英文摘要
Background: Iterative feedback loops are the dominant paradigm for improving LLM-generated Infrastructure-as-Code (IaC): validators such as Checkov and terraform validate feed error signals back for successive repair attempts. Prior work reports cumulative-best metrics, which are non-decreasing by construction, so the raw per-iteration security trajectory has never been examined for IaC. Aims: We study security regression (a previously-passing CIS Benchmark check that fails after a repair iteration) to determine whether and how often iterative LLM repair degrades security while fixing other issues. Method: We analyze 5,968 scenario timelines from the IaC-Eval benchmark, each one scenario run through one configuration for up to 5 repair iterations. The 15 configurations (six model-specific RAG, nine model-aggregated non-RAG, three temperatures each) yield 4,440 iteration transitions with Checkov data on both sides. We track 30 individual CIS check IDs and classify root causes from code diffs, under two detection modes: standard (inclusive) and strict (exclusive check failures only). Results: Under standard detection, 13.8% of scenarios (24.8% of transitions) exhibit at least one regression. Under strict detection the rate falls to 3.3% of scenarios (5.2% of transitions), indicating most apparent regressions are multi-resource measurement artifacts. Resource restructuring (79.0%) is the dominant root cause. Regression transitions show 2.6x more code churn (Cohen's d=0.90) and 4.9x higher strict-mode check volatility (d=1.49). Of standard-mode regressions, 36.6% self-correct within an average of 1.2 iterations; iteration 3 is the optimal stopping point. Conclusions: Iterative IaC repair does introduce security regressions, but the conservative, defensible rate is about 3.3% of scenarios. Our findings motivate security-aware feedback-loop design and actionable iteration-budget guidance.
Comments20 pages, 3 figures, 5 tables. Accepted at the 20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026). To appear in LIPIcs Vol. 394. v2: corrected the bibliographic record of one reference (preprint, not a journal article) and added the related-version link to the published LIPIcs article