arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

超越代码生成:智能体软件开发生命周期中的可靠性、验证与成本经济学

Beyond Code Generation: Reliability, Verification, and Cost Economics in the Agentic Software Development Lifecycle

Happy Bhati

arXiv 2609.04681首次发表:更新:

AI 中文总结

本文综合多类研究证据,提出智能体SDLC相关工程概念,探讨AI编码智能体在软件开发全流程的可靠性、成本等问题,核心转向生产合格价值的量化评估。

AI 中文摘要

AI编码系统正从自动补全和聊天功能转向可检查代码仓库、编辑多个文件、运行工具、编写测试、发起拉取请求并在有限监督下长期工作的智能体。这种能力改变了软件交付的瓶颈。近期实地研究显示编码活动有显著提升,但新证据也表明,这些提升在编写代码与交付可靠软件之间急剧减弱。评审、集成、测试、安全、部署及生产操作仍是制约阶段,而成本模式正从可预测的每席位许可转向可变的token、工具、沙箱、CI及返工成本。本文综合了2024年至2026年9月发布的同行评审软件工程研究、大学研究、基准审计、大型科技公司生产报告、开发者遥测数据及成本管理证据,未声称开展新的模型实验,数值发现均归因于原始研究。该综合研究提出四个工程概念:智能体SDLC吞吐量悖论、生产合格变更(PQC)、验证税,以及受成本、可靠性和人力注意力预算约束分配自主权的智能体SDLC控制平面。随后,基于证据的视野将当前受监督智能体映射到未来受政策约束的软件工厂,核心研究问题从智能体能生成多少代码转变为工程系统每美元、每评审员小时及每单位操作风险能交付多少生产合格价值。

英文摘要

AI coding systems are moving from autocomplete and chat toward agents that can inspect repositories, edit multiple files, run tools, write tests, open pull requests, and work for long periods with limited supervision. This capability changes the bottleneck in software delivery. Recent field studies show meaningful gains in coding activity, but newer evidence also shows that those gains attenuate sharply between writing code and shipping reliable software. Review, integration, testing, security, deployment, and production operations remain constraining stages, while the economics are shifting from predictable per-seat licensing toward variable token, tool, sandbox, CI, and rework costs. This paper synthesizes peer-reviewed software-engineering research, university studies, benchmark audits, production reports from major technology companies, developer telemetry, and cost-management evidence released primarily from 2024 through September 2026. No new model experiment is claimed; numerical findings remain attributed to their original studies. The synthesis proposes four engineering concepts: the Agentic SDLC Throughput Paradox, Production-Qualified Change (PQC), the Verification Tax, and an Agentic SDLC Control Plane that allocates autonomy subject to cost, reliability, and human-attention budgets. An evidence-based horizon then maps today's supervised agents to future policy-bounded software factories. The central research question shifts from how much code an agent can generate to how much production-qualified value an engineering system can deliver per dollar, per reviewer-hour, and per unit of operational risk.

Comments18 pages, 6 figures, 3 tables. Systems synthesis and research agenda on agentic software engineering, code review, testing, reliability, and AI cost. No new experimental measurements are claimed; empirical and company-reported results are attributed to the cited sources

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑