AI 中文总结
本研究通过模糊测试对比AI生成与人工编写的Linux工具,发现AI代码可靠性相当或更优,但需人工监督,且工作流可作规范以提升可持续性。
AI 中文摘要
基于智能体AI的软件开发有望实现更快的软件完成速度、更高的程序员效率以及更可靠的代码。问题在于,我们如何客观地验证这些说法?在本项目中,我们尝试基于三种实践来回答这一问题。首先,我们应用了典型的智能体AI最佳实践工作流进行软件开发。其次,我们的目标程序是十个知名的、达到发布质量的人工编写的Linux实用程序,以便将AI生成的代码与具体的基准真相进行比较。第三,我们将可靠性度量建立在广泛使用的测试技术——模糊随机测试之上。在此测试中,我们既使用了经典的黑盒、生成式测试,也使用了更现代的、基于覆盖率引导的(灰盒、变异式)测试,采用AFL++工具。我们发现,AI生成的实用程序版本通常与这些程序最新人工生成版本一样可靠,甚至往往更可靠。虽然AI生成的版本确实存在一些失败,但其发生频率低于标准代码库中的代码。有趣的是,AI生成的代码不太可能出现内存错误(如缓冲区溢出)等失败,但更可能出现挂起,如无限循环。此外,我们验证了使用智能体AI生成健壮且可靠的软件需要谨慎的实践和人工监督。代码质量高度依赖于所使用的提示词和技能,以及指导该过程的人类如何响应。我们还证明了,使用智能体AI工作流进行软件开发(及其提示词和技能)可以成为代码的规范,从而带来软件的成本效益可持续性。
英文摘要
Agentic-AI based software development offers the promise of faster completion of the software, greater programmer efficiency, and more reliable code. The question is how can we verify these claims in an objective way? In this project, we attempted to answer this question based on three practices. First, we applied a typical best-practices agentic AI workflow for software development. Second, our target programs were ten well-known, release-quality human-written Linux utility programs so that we could compare the AI-generated code against a concrete ground truth. Third, we based our measure of reliability on a widely used testing technique, fuzz random testing. For this testing, we used both classic black box, generational testing and more modern coverage guided (gray box, mutational) testing using AFL++. We found that the AI-generated versions of the utility programs were typically as reliable - often more reliable - than the latest human-generated versions of these programs. While the AI-generated versions did have some failures, they were less common than the code from the standard repositories. Interestingly, the AI-generated code was less likely to have failures such as memory errors (such as buffer overflows) but more likely to have hangs such as infinite loops. In addition, we verified that generating robust and reliable software using agentic AI requires careful practice and human supervision. The quality of the code is highly dependent on the prompts and skills used, and how the human directing the process responds. We also demonstrated that using agentic AI workflow for software development (with its prompts and skills) can become a specification of the code that leads to cost-effective sustainability of the software.