AI 中文总结
本文提出自动化多维度评估框架,针对C#代码生成评估GPT等四个LLM,发现Pass@k排名无法反映LLM在软件工程场景的完整性能概况,还明确了GPT的双峰失败行为。
AI 中文摘要
评估大型语言模型(LLM)的代码生成质量,不仅需检查生成代码是否正确,还需检查其可维护性、效率和风格合理性,这些都是软件工程从业者直接关注的品质。现有基准将评估简化为单一的Pass@k指标,这掩盖了功能正确性与结构质量之间的关键权衡;此外,现有基准几乎只关注Python,未针对C#和.NET等与企业相关的生态系统开展专门评估。本文提出了一种用于C#代码生成的自动化多维度评估框架,将其应用于四个最先进的LLM:GPT、Gemini、Claude和Grok。我们基于HumanEval衍生的85个算法任务开展受控实验,共生成并评估340个解决方案,每个解决方案在三个独立维度接受评估:通过自动化单元测试评估功能正确性、通过Roslyn AST分析评估静态代码质量、通过对抗性BenchmarkDotNet分析评估运行时效率。我们的核心发现表明,正确性与质量属性之间存在显著差距(皮尔逊相关系数r=0.075),证明Pass@k排名会系统性地歪曲软件工程场景中LLM的完整性能概况;我们还明确了GPT的双峰失败行为。
英文摘要
Evaluating Large Language Model (LLM) code generation quality requires examining not just whether the generated code is correct, but whether it is maintainable, efficient, and stylistically sound, all of which are qualities of direct importance to software engineering practitioners. Existing benchmarks reduce evaluation to a single Pass@k metric, which obscures critical trade-offs between functional correctness and structural quality. A further limitation is the near-exclusive focus on Python, leaving enterprise-relevant ecosystems such as C# and .NET without dedicated evaluation. This paper presents an automated, multi-dimensional evaluation framework for C# code generation, applying it to four state-of-the-art LLMs: GPT, Gemini, Claude, and Grok. We conduct a controlled experiment across 85 algorithmic tasks derived from HumanEval, generating and evaluating 340 solutions in total, in which each solution is assessed across three independent dimensions: functional correctness via automated unit testing, static code quality via Roslyn AST analysis, and runtime efficiency via adversarial BenchmarkDotNet profiling. Our central finding reveals a substantial gap between correctness and quality attributes (Pearson r = 0.075), demonstrating that Pass@k rankings systematically misrepresent the full LLM performance profile in software engineering contexts. We further characterize GPT's bimodal failure behavior.
Comments12 pages, 8 figures. Accepted at the Second Workshop on Large Language Models for Generative Software Engineering (LLM4SE 2026), co-located with STAF 2026, Rennes, France, June 29 - July 3, 2026