发表机构
IReSCoMath Research Laboratory, Faculty of Sciences, University of Gabes; National School of Engineering, Gabes; Department of Electrical & Software Engineering, University of Calgary(加贝斯大学理学院 IReSCoMath 研究实验室; 加贝斯国立工程学院; 卡尔加里大学电气与软件工程系)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对LLM生成代码中的库使用错误,提出一种智能体方法,通过任务分析、文档接地、代码生成和自动验证,将库相关错误减少38.1%-54.6%,代码正确性提升16%。
AI 中文摘要
软件从业者越来越依赖大语言模型(LLM)来生成集成外部库的代码。然而,LLM 经常产生错误的库使用,例如无效的导入、过时的 API 调用和幻觉依赖,导致编译或运行时失败,从而降低了 AI 辅助软件开发的可靠性。在本文中,我们提出了一种智能体方法来缓解 LLM 生成代码中的库相关错误。更具体地说,我们首先进行了一项探索性研究,以表征 LLM 产生的库相关问题。我们对 100 个 LLM 生成的代码文件的分析显示,84% 的生成文件包含至少一个库相关错误,重复出现的模式包括错误的导入路径、缺失的导入、幻觉库、过时的库使用和未使用的导入。基于这些发现,我们设计了一种智能体方法,该方法整合了任务分析、文档接地、代码生成和自动验证,以在代码合成过程中改进库的使用。我们在从快速发展的 Python 框架(包括 LangChain 和 AutoGen)的真实世界实现中衍生的 300 个代码生成任务上评估了我们的方法,涉及五个 LLM:GPT-5、DeepSeek-V3、Qwen3、Mistral 和 Llama 3。结果表明,我们的方法在所有评估的模型中持续提高了代码生成质量,将库相关错误减少了 38.1% - 54.6%,并将代码正确性提高了最多 16%。
英文摘要
Software practitioners increasingly rely on Large Language Models (LLMs) to generate code that integrates external libraries. However, LLMs often produce incorrect library usage, such as invalid imports, outdated API calls, and hallucinated dependencies, leading to compilation or runtime failures that reduce the reliability of AI-assisted software development. In this paper, we propose an agentic approach to mitigate libraryrelated errors in LLM-generated code. More specifically, we first conduct an exploratory study to characterize the library-related issues produced by LLMs. Our analysis of 100 LLM-generated code files reveals that 84% of generated files contain at least one library-related error, with recurring patterns including incorrect import paths, missing imports, hallucinated libraries, deprecated library usage, and unused imports. Based on these findings, we design an agentic approach that integrates task analysis, documentation grounding, code generation, and automated validation to improve library usage during code synthesis. We evaluate our approach on 300 code generation tasks derived from realworld implementations of rapidly evolving Python frameworks, including LangChain and AutoGen, across five LLMs: GPT-5, DeepSeek-V3, Qwen3, Mistral, and Llama 3. The results show that our approach consistently improves code generation quality across all evaluated models, reducing library-related errors by 38.1% - 54.6% and increasing code correctness by up to 16%.