发表机构
The Pennsylvania State University(宾夕法尼亚州立大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究大语言模型生成C代码可靠性问题,提出结合编译诊断、CodeQL静态分析等及检索先前修复模式的分析与修复工作流程,在相关C编程任务上评估,该方法有效降低了基线模型的编译失败率和安全缺陷率,提升了代码安全性。
AI 中文摘要
大语言模型能从自然语言描述生成C代码,但生成的程序常含安全漏洞和编译错误,对嵌入式及资源受限系统构成风险。本文研究反馈和检索如何提高大语言模型生成的C代码的可靠性。提出一种分析与修复工作流程,结合编译诊断、CodeQL静态分析、KLEE符号执行及检索先前修复模式进行迭代优化。在5000个涉及嵌入式相关漏洞的C编程任务上评估,基线模型存在可靠性差距,编译失败率高达46%,安全缺陷率高达49%。我们的方法改善了这两个指标。对于CodeLlama 7B,安全缺陷率从49%降至19%,CodeQL总错误从15088降至2463(83.7%)。对于DeepSeek Coder 1.3B,编译失败从42%降至22%,安全缺陷从35%降至15%。结果表明集成轻量级分析工具可提高大语言模型生成的代码在嵌入式开发中的安全性。
英文摘要
Large language models can generate C code from natural-language descriptions, but resulting programs often contain security vulnerabilities and compilation errors, posing risks for embedded and resource-constrained systems. This work investigates how feedback and retrieval improve reliability of LLM-generated C code. We present an analysis-and-repair workflow that combines compilation diagnostics, CodeQL static analysis, and KLEE symbolic execution with retrieval of prior repair patterns for iterative refinement. Evaluated on 5,000 C programming tasks exercising embedded relevant vulnerabilities, baseline models show substantial reliability gaps, with compilation failure rates up to 46% and security defect rates up to 49%. Our approach improves both metrics. For CodeLlama 7B, security defect rates decrease from 49% to 19% and total CodeQL errors drop from 15,088 to 2,463 (83.7%). For DeepSeek Coder 1.3B, compilation failures are reduced from 42% to 22% and security defects from 35% to 15%. These results show that integrating lightweight analysis tools can improve the safety of LLM-generated code for embedded development.