PolyCodeEval:从函数到代码仓库的多语言代码生成基准测试
PolyCodeEval: Benchmarking Multilingual Code Generation from Functions to Repositories
查看机构详情
- Tianjin University(天津大学)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
PolyCodeEval是覆盖5种语言58个开源仓库的多粒度代码生成基准,评估前沿模型等发现现有方法生成正确函数等的比例有限,且性能随语言和粒度波动,相关函数上下文可提升函数生成正确性。
中文摘要 AI 辅助
随着大语言模型日益向代码仓库级软件工程方向发展,现有的代码生成基准在语言覆盖范围、任务粒度和评估协议方面仍存在碎片化问题,阻碍了系统性比较。为解决这一缺口,我们提出PolyCodeEval,一个统一的多语言、多粒度代码生成基准。它包含2590个代码生成任务,覆盖从函数到代码仓库的范围,源自5个编程语言的58个真实可执行开源代码仓库。所有任务均采用统一的基于执行的协议进行评估,且针对不同生成目标定制了集成流程。基于该基准,我们评估了前沿大语言模型、最先进的专用方法以及通用编码智能体。结果显示,现有方法在跨粒度和语言生成完整代码片段方面仍存在困难。具体而言,所研究的方法生成的正确函数、文件、代码仓库的比例分别最高为71.7%、76.7%和31.0%,且性能随语言差异大幅波动。配对实验进一步表明,同一文件中相关函数的实现上下文可提升函数生成的可执行正确性。方法排名也随任务粒度和编程语言而变化,凸显了多语言、多粒度评估对全面评估代码生成能力的重要性。
英文摘要
As large language models increasingly move toward repository-level software engineering, existing code-generation benchmarks remain fragmented across language coverage, task granularity, and evaluation protocols, impeding systematic comparison. To address this gap, we present PolyCodeEval, a unified multilingual and multi-granularity benchmark for code generation. It comprises 2,590 code generation tasks spanning functions to repositories, derived from 58 real, executable open-source repositories in five programming languages. All tasks are evaluated under a unified execution-based protocol with integration procedures tailored to their generation targets. Building on this benchmark, we evaluate frontier large language models, state-of-the-art specialized methods, and general coding agents. Our results show that existing approaches still struggle to correctly generate complete code fragments across granularities and languages. Specifically, the studied methods generate at most 71.7%, 76.7%, and 31.0% correct functions, files, and repositories, respectively, with performance varying widely across languages. Paired experiments further show that implementation context from related functions in the same file improves the executable correctness of function generation. Method rankings also vary across task granularities and programming languages, highlighting the importance of multilingual, multi-granularity evaluation for comprehensively assessing code generation capabilities.