发表机构
Tampere University; Stony Brook University(坦佩雷大学; 石溪大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出CodeAssay基准,涵盖185个Python任务,结合多类测试与度量,经审计后发现模型正确性标签有9.0%变化,且验证了多类LLM代码生成的评估特性。
AI 中文摘要
目前,评估大语言模型(LLM)代码生成能力时,越来越多地采用基于测试的基准,此类评估的有效性依赖于其参考基准和测试的可靠性,而基于测试的正确性仅能捕捉生成代码可观测属性的一部分。本文提出CodeAssay,这是一个按分类法构建的基准,涵盖10个软件工程类别的185个Python任务,它结合了经审计的基准真值、用于生成与修复的公开测试、用于评分的隐藏测试、基于变异的测试套件验证以及选定的代码属性度量。在审计后对模型修复输出重新评分,1890个正确性标签中有170个发生变化(占比9.0%),且最佳到最差模型的测量差距从11.9个百分点扩大至23.7个百分点,不过总体正确性几乎保持不变。完整测试套件和隐藏测试套件的变异分数分别达到82.6%和74.8%。在7个专有LLM中,标准提示下的正确性范围为77.3%至98.9%,21组模型对中有12组存在显著差异。在14种模型-提示配置均能解决的120个任务中,没有模型在所有选定代码属性上表现最佳。面向安全的提示未使正确性发生显著变化,也未持续减少选定的静态分析结果,但会增加所有模型的程序长度和圈复杂度。这些发现表明,对LLM生成代码的可靠评估需要经过验证的基准真值、受保护的测试以及多个明确解读的度量,CodeAssay为AI增强型软件开发中基于证据的模型评估提供了可复现的基础。
英文摘要
Large Language Models are increasingly evaluated for code generation using test-based benchmarks. The validity of such evaluations depends on the reliability of their references and tests, while test-based correctness captures only part of the observable properties of generated code. We present CodeAssay, a taxonomy-first benchmark of 185 Python tasks across ten software-engineering categories. It combines audited ground truth, public tests for generation and repair, hidden tests for grading, mutation-based test-suite validation, and selected code-property measures. Regrading fixed model outputs after the audit changed 170 of 1,890 correctness labels (9.0%) and increased the measured best-to-worst model spread from 11.9 to 23.7 percentage points, although aggregate correctness remained nearly unchanged. The complete and hidden test suites achieved mutation scores of 82.6% and 74.8%, respectively. Across seven proprietary LLMs, standard-prompt correctness ranged from 77.3% to 98.9%, with significant differences in 12 of 21 model pairs. On the 120 tasks solved by all 14 model-prompt configurations, no model performed best across all selected code properties. A security-focused prompt produced no significant change in correctness or consistent reduction in the selected static-analysis findings, while increasing program length and cyclomatic complexity across all models. These findings show that reliable evaluation of LLM-generated code requires validated ground truth, protected tests, and multiple explicitly interpreted measures. CodeAssay provides a reproducible basis for evidence-based model evaluation in AI-augmented software development.
CommentsPROFES 2026: 27th International Conference on Product-Focused Software Process Improvement