arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.32414cs.SE

AI如何改变DevOps性能:一种基于机制的模拟

How AI Changes DevOps Performance: A Mechanism-Based Simulation

Mamdouh Alenezi

首次发表
浏览论文内容

中文总结 AI 辅助

本研究通过离散事件模拟,揭示AI生成代码与智能体式AI对DevOps性能指标的不同影响机制,提出测量规则与验证协议。

中文摘要 AI 辅助

背景。用于软件开发的AI工具已从代码补全发展到自主智能体,但关于其对DORA性能影响的证据混杂,大多为横截面研究,且常常将AI生成代码(AGC)与智能体式AI(AGT)混为一谈。目标。我们考察AGC和AGT分别及联合对部署频率、前置时间、变更失败率、恢复时间和部署返工率的影响。方法。我们审查证据,形式化一个包含四个能力调节变量的双构念模型,并实现交付管道的离散事件模拟。六项实验包括因子设计、Shapley分解、600次参数抽取的敏感性分析、固定需求稳健性检验,以及包含反事实的5000个模拟团队的错开采用面板。结果。在两种能力概况下,AGC均增加了失败率、返工和恢复时间。当AI增加变更量时,AGC仅在审查能力有闲置时缩短前置时间(-13%),一旦审查饱和则延长前置时间;在低能力团队中,其部署频率的提升完全来自计划外返工。AGT改善了部署频率(+31%至+34%)、前置时间(-8%至-22%)和恢复时间(-29%至-38%),但其稳定性效应依赖于自主CI修复掩盖缺陷,且在高能力团队中逃逸缺陷增加。能力降低了AGC的绝对失败率惩罚(7.9个百分点对比1.5个百分点),但未降低其相对惩罚(+34%对比+84%)。双向固定效应低估了采用效应8-16%;检测所建模的稳定性效应需要约50个团队,其能力调节效应需要约400个团队。结论。在该模型内,DORA指标通过排队、批处理和监督对AI作出响应,而不仅仅通过代码质量。我们推导出测量规则、智能体式修复的护栏,以及基于样本量的现场验证协议。

英文摘要

Context. AI tools for software development have progressed from code completion to autonomous agents, yet evidence on their effects on DORA performance is mixed, mostly cross-sectional, and often conflates AI-generated code (AGC) with agentic AI (AGT). Objective. We examine how AGC and AGT, separately and jointly, affect deployment frequency, lead time, change failure rate, recovery time, and deployment rework rate. Method. We audit the evidence, formalise a two-construct model with four capability moderators, and implement a discrete-event simulation of a delivery pipeline. Six experiments include a factorial design, Shapley decompositions, sensitivity analysis over 600 parameter draws, a fixed-demand robustness test, and a staggered-adoption panel of 5,000 simulated teams with counterfactuals. Results. AGC increased failure rate, rework, and recovery time in both capability profiles. When AI increased change volume, AGC shortened lead time only where review capacity was spare (-13%) and lengthened it once review saturated; in the low-capability team, its deployment-frequency gain came entirely from unplanned rework. AGT improved deployment frequency (+31% to +34%), lead time (-8% to -22%), and recovery time (-29% to -38%), but its stability effect depended on autonomous CI repair masking defects, and escaped defects increased in the high-capability team. Capability reduced AGC's absolute failure-rate penalty (7.9 vs. 1.5 points) but not its relative penalty (+34% vs. +84%). Two-way fixed effects underestimated adoption effects by 8-16%; detecting the modelled stability effect required about 50 teams, and its capability moderation about 400. Conclusions. Within the model, DORA metrics respond to AI through queueing, batching, and oversight, not only code quality. We derive measurement rules, guardrails for agentic remediation, and a sample-size-informed field-validation protocol.

发表机构

  • SDAIA Academy, Saudi Data and Artificial Intelligence Authority (SDAIA)(沙特数据和人工智能管理局(SDAIA)学院)

机构由 AI 辅助整理,请以论文原文为准。

↑