arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

PROBE:大语言模型中代码生成的基准测试

PROBE: Benchmarking Code Generation in Large Language Models

Rodrigo Pato Nogueira, Marco Vieira, João R. Campos

arXiv 2607.13820首次发表:更新:

AI 中文总结

针对大语言模型代码生成评估不足,介绍PROBE基准框架,基于多维度评估代码,用于评估多种模型和语言,发现LLMs虽有成果,但处理难题及小模型在特定语言上有困难,且常因易避免错误失败。

AI 中文摘要

大语言模型(LLMs)在日常软件工程任务尤其是自动代码生成中使用日益广泛,但仍远非完美,系统公正评估很关键。现有代码生成基准测试有局限。为此引入PROBE,一个可扩展基准框架,基于多样明确指标、代表性工作负载、不同提示模板和稳健实验程序构建系统结构。通过三个互补维度评估LLMs生成的代码,用其评估多种模型和语言。研究发现LLMs虽有成果,但处理难题及小模型在资源少的语言上有困难,还常因基本易避免错误失败,凸显自动生成代码不可靠。

英文摘要

Large Language Models (LLMs) are increasingly being used in everyday software engineering tasks, particularly in automated code generation. Despite their widespread adoption, these models remain far from perfect, making systematic and fair evaluation essential to understand their strengths and limitations. In the context of code generation, existing benchmarks are limited: they often target a single programming language and rely primarily on unit test outcomes, while overlooking other critical dimensions such as the overall quality of the generated code and its closeness to a valid solution. To address these gaps, we introduce PROBE, an extensible benchmark framework that, unlike prior work, establishes a systematic structure built on diverse and well-defined metrics, representative workloads, varied prompt templates, and a robust experimental procedure. In practice, the code generated by the LLMs is evaluated along three complementary dimensions: functional correctness, proximity to valid solutions, and code quality, enabling a comprehensive assessment of performance. We use PROBE to evaluate four open-source and two proprietary models under three prompting strategies across five programming languages. We further complement this analysis with a study of common errors in the code and provide concrete examples, offering clearer insight into where LLMs tend to struggle. Our findings show that, while LLMs achieve promising results, they struggle with harder problems and, in the case of smaller models, with programming languages that have fewer available resources for training, and they often fail due to fundamental and easily avoidable errors that underscore the unreliability of automatically generated code.

CommentsAccepted for publication in Empirical Software Engineering (Springer)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑