AI 中文总结
本文通过构建DOC2CI基准评估LLM生成CI配置的能力,发现其结构有效性不足,提出无需训练的模式引导修复方法提升有效性,为相关工具开发提供依据。
AI 中文摘要
采用持续集成(Continuous Integration,CI)通常需要编写易出错且难以维护的YAML配置。尽管大语言模型(Large Language Models,LLM)在软件工程中的应用日益增多,但它们跨服务和模型家族从自然语言生成CI配置的能力仍不明确。本文开展了一项关于使用LLM生成CI配置的大型实证研究。我们推出了DOC2CI,这是一个从4个CI服务的官方文档中收集的3363组描述- YAML对的基准,评估了7B至34B参数范围内的14个开源权重模型,以及GPT-4o和GPT-4.1,生成了超过53000个配置。我们评估了参考对齐和模式有效性,以确定生成的配置在结构上是否有效。我们还通过对385个配置的手动分析开发了故障分类法,并研究了LLM产生差异的原因。在所有模型和服务中,精确的参考复现率从未超过3.1%,虽然97%的输出可解析为YAML,但仅有71%符合服务模式。更大规模的模型提升了结构有效性,但代码专业化相较于同类通用模型并未提供一致的优势。模型差异主要由输出完整性驱动:对于同一请求,部分模型生成预期片段,而其他模型生成完整工作流。最后,一种无需训练的模式引导修复方法将模式有效性提升至94%,而微调提升了与文档的相似性但降低了独立有效性。这表明相似性和有效性是CI生成的不同目标,推动了面向基于LLM的配置生成的感知模式评估与工具开发。
英文摘要
Adopting Continuous Integration (CI) often requires writing YAML configurations that are error-prone and challenging to maintain. Despite increasing LLM use in software engineering, their ability to generate CI configurations from natural language across services and model families remains unclear. This paper presents a large empirical study on using LLMs to generate CI configurations. We introduce DOC2CI, a benchmark of 3,363 description-to-YAML pairs collected from the official documentation of four CI services, and evaluate 14 open-weight models from 7B-34B parameters together with GPT-4o and GPT-4.1, producing over 53,000 configurations. We assess both reference alignment and schema validity to determine whether the generated configurations are structurally valid. We further develop a failure taxonomy from a manual analysis of 385 configurations and examine why LLMs disagree. Across models and services, exact reference reproduction never exceeds 3.1%, and while 97% of outputs parse as YAML, only 71% satisfy service schemas. Larger models improve structural validity, but code specialization provides no consistent advantage over comparable general models. Model differences are driven largely by output completeness: for the same request, some models generate the expected fragment while others produce a full workflow. Finally, a training-free schema-guided repair method improves schema validity to 94%, while fine-tuning improves similarity to documentation but reduces standalone validity. This suggests that similarity and validity are distinct objectives for CI generation and motivate schema-aware evaluation and tooling for LLM-based configuration generation.