实践中不可靠?对大语言模型生成代码中的错误的综合研究
Unreliable in Practice? A Comprehensive Study of Errors in LLM-Generated Code
浏览论文内容
中文总结 AI 辅助
本研究分析7款LLMs生成的86726个含错代码样本,对比不同模型、语言的错误模式,发现大型模型也常犯简单错误,代码常省略关键安全检查,为提升LLM生成代码可靠性提供依据。
中文摘要 AI 辅助
大语言模型(LLMs)被广泛用于编程,有报告显示AI如今生成的生产代码占比越来越大。研究表明,LLMs能显著提升开发者的生产力,但在更复杂的编程任务上仍存在困难。正如理解人类编写代码的错误模式对提升软件质量至关重要一样,识别和表征LLMs生成代码中的错误,对于设定合理预期、设计缓解策略也十分关键。现有研究范围有限,通常仅聚焦单一编程语言、少量问题或有限的模型选择,因此目前仍未全面明确哪些错误是常见的,哪些是特定模型或语言特有的。为填补这些空白、更深入地理解LLMs生成代码的质量,我们分析了包含编译或运行时错误的86726个代码样本,这些样本由7款LLMs生成,涉及4种编译型语言。我们利用一款LLM按根本原因对错误进行分类,对这些分类进行人工验证后开展对比分析,再用这些标注数据衡量不同模型、语言及问题难度下的错误发生率,以识别常见错误模式。结果显示,尽管错误类型在不同语言和模型间差异显著,但即使是最大型的模型也常犯简单错误;我们还发现,生成的代码常省略基础的输入验证或内存安全检查,这可能导致溢出、资源耗尽或其他可靠性/安全问题。
英文摘要
Large Language Models (LLMs) are being widely used for coding, with reports indicating that AI now generates an increasing share of production code. Studies show that LLMs can significantly improve developer productivity, yet they still struggle with more complex coding tasks. Just as understanding error modes in human-written code has been central to improving software quality, identifying and characterizing the errors in LLM-generated code is critical for setting realistic expectations and designing mitigation strategies. Prior research has been limited in scope, often focusing on a single language, a small number of problems, or a limited selection of models. As a result, there is still no comprehensive understanding of which errors are common and which are specific to certain models or languages. To address these gaps and develop a deeper understanding of the quality of LLM-generated code, we analyzed a corpus of 86,726 code samples that contained compilation or runtime errors. These samples were generated by seven LLMs across four compiled languages. We classified errors by their underlying causes using an LLM, manually validated these classifications, and performed a comparative analysis. This labeled data is then used to measure error prevalence by model, language, and problem difficulty, to identify common error patterns. Results show that, although error types vary strongly across languages and models, even the largest models frequently make simple mistakes. We also observe that generated code often omits basic input validation or memory-safety checks, which can lead to overflows, resource exhaustion, or other reliability/security issues.