我与你的区别是什么?基准测试人类编写与AI生成代码之间的质量差距
What is the Difference Between Me and You? Benchmarking the Quality Gap Between Human-Written and AI-Generated Code
- University of Naples Federico II(那不勒斯费德里科二世大学)
- University of North Carolina at Charlotte(北卡罗来纳大学夏洛特分校)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本研究通过大规模函数对比较人类与AI生成代码,发现AI代码更紧凑模板化,缺陷类型不同,安全性因语言而异,并发布CQBench基准。
AI中文摘要:
AI编码助手正逐渐成为生产软件的合著者,然而对其评估主要集中于功能正确性,这遗留了一个问题:其生成的代码是否在主导生命周期成本的质量维度上区别于人类代码。我们大规模比较了人类编写与AI生成的代码:涵盖Python、Java和C语言的787,562个函数对,每个从开源仓库挖掘的人类函数与其文档字符串生成的三款AI助手(OpenAI GPT模型、DeepSeek-Coder、Qwen2.5-Coder)的实现配对。我们刻画了结构复杂性和统计自然度,并将静态分析发现映射到正交缺陷分类(用于缺陷)和常见弱点枚举(用于漏洞),使作者和语言可直接比较。AI生成的代码在结构上更紧凑且风格模板化:其大小和分支约为人类代码的一半,在风格层面聚类分离。缺陷概况在种类上有所不同:人类代码集中于成熟代码库的问题,AI代码则表现为重复的样板代码;安全性依赖语言,LLM在Python和Java中产生更多且更严重的发现,但在C语言中产生的高严重性内存安全发现少于人类。一旦控制大小,复杂性指标携带的信号很少,而自然度能区分作者。最后,我们发布了CQBench,一个包含27,346个易出问题任务、基线及用于质量保证和安全测试评估流程的基准。
英文摘要:
AI coding assistants are becoming co-authors of production software, yet their evaluation centers on functional correctness, leaving open whether their code differs from human code in the quality dimensions dominating lifecycle cost. We compare human-written and AI-generated code at scale: 787,562 function pairs across Python, Java, and C, each human function mined from open-source repositories paired with implementations generated from its docstring by three AI assistants (OpenAI GPT models, DeepSeek-Coder, Qwen2.5-Coder). We characterize structural complexity and statistical naturalness, and map static-analysis findings onto Orthogonal Defect Classification for defects and the Common Weakness Enumeration for vulnerabilities, making authors and languages directly comparable. AI-generated code is structurally compressed and stylistically templated: roughly half the size and branching of human code, clustering apart at the style level. Defect profiles differ in kind: human code concentrates issues of mature codebases, AI code repetitive boilerplate; security is language-dependent, with LLMs producing more, and more severe, findings in Python and Java but fewer high-severity memory-safety findings than humans in C. Once size is controlled for, complexity metrics carry little signal, while naturalness separates authors. Finally, we release CQBench, a benchmark of 27,346 issue-prone tasks with baselines and an evaluation pipeline for quality assurance and security testing.