arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.20478cs.SEcs.AI

用于基础设施即代码生成的智能语言模型的验证优先评估

Verifier-First Evaluation of Agentic LLMs for Infrastructure-as-Code Generation

Mohamed Jouini

首次发表
浏览论文内容

中文总结 AI 辅助

研究针对自然语言生成基础设施即代码的问题,对七种智能策略在特定基准测试上进行验证优先评估,通过多种方法得出如主动检索提升性能、迭代优化有收敛效果等五个主要发现,为相关研究提供了重要参考。

中文摘要 AI 辅助

从自然语言生成基础设施即代码(IaC)需要满足提供商模式、依赖规划和组织政策约束,而不仅仅是生成语法上合理的配置。我们对七种用于Terraform生成的智能策略进行了验证优先的实证研究,这些策略在IaC-Eval v2上进行评估,这是一个具有Rego v1意图策略的现代化186任务AWS/Terraform基准测试。我们的评估将失败分为三个验证阶段(terraform验证、terraform计划、opa评估),并在所有成对比较(n = 186,α = 0.05)上应用带有威尔逊置信区间的麦克尼马尔检验。我们报告了五个主要发现。(1)通过带有MCP或ChromaDB支持的RAG的ReAct智能体进行主动检索,将Qwen2.5-Coder 7B的pass@1从14.0%提高到45.7%(p < 0.0001),主要是通过将VALIDATE_FAIL从144个任务减少到66个任务。(2)通过验证器反馈进行迭代优化,Qwen 7B达到62.9%,GPT-4o达到84.4%的pass@1,呈现出二元收敛——任务要么在一次重试中解决,要么耗尽预算。(3)GEPA反射式指令优化仅使用80次验证器引导的展开,就将主动RAG基线提高了7.5个百分点(p = 0.026),这表明提示优化器可以在不更新权重的情况下改善可验证的IaC生成。(4)SIMBA无教师示范注入在没有检索基础设施的情况下实现了与主动RAG相当的性能(p = 1.0),但未能解决占主导地位的自定义属性错误类(50%的失败)。(5)一项诊断性Rego注入实验表明,当策略文本可见时,79%的细化后OPA失败是可解决的信息差距失败(p = 0.016),这激发了政策制定。

英文摘要

Infrastructure-as-Code (IaC) generation from natural language requires satisfying provider schemas, dependency planning, and organizational policy constraints, not merely producing syntactically plausible configurations. We present a verifier-first empirical study of seven agentic strategies for Terraform generation evaluated on IaC-Eval v2, a modernized 186-task AWS/Terraform benchmark with Rego v1 intent policies. Our evaluation separates failures into three verifier stages (terraform validate, terraform plan, opa eval) and applies McNemar's test with Wilson confidence intervals on all pairwise comparisons (n=186, alpha=0.05). We report five principal findings. (1) Active retrieval via ReAct agents with MCP or ChromaDB-backed RAG raises Qwen2.5-Coder 7B from 14.0% to 45.7% pass@1 (p<0.0001), primarily by reducing VALIDATE_FAIL from 144 to 66 tasks. (2) Iterative refinement with verifier feedback achieves 62.9% (Qwen 7B) and 84.4% (GPT-4o) pass@1, exhibiting binary convergence -- tasks either resolve in one retry or exhaust the budget. (3) GEPA reflective instruction optimization raises the Active RAG baseline by +7.5 pp (p=0.026) using only 80 verifier-guided rollouts, providing evidence that prompt optimizers can improve verifiable IaC generation without weight updates. (4) SIMBA teacher-free demonstration injection achieves performance equivalent to Active RAG (p=1.0) without retrieval infrastructure, but fails to address the dominant SELF_DEFINED_PROPERTY error class (50% of failures). (5) A diagnostic Rego-injection experiment shows that 79% of post-refinement OPA failures are information-gap failures resolvable when policy text is visible (p=0.016), motivating policy

补充信息

↑