发表机构
Singapore Management University; Nanjing University; SUN YAT-SEN University; Home Team Science & Technology Agency(新加坡管理大学; 南京大学; 孙中山大学; 家庭团队科技局)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究大型语言模型处理科学代码的能力,提出SciCodePile语料库及可执行基准测试,评估15个模型在三项任务上的表现,发现科学代码生成极具挑战,同时展示了该语料库在训练上的效用,代码和数据可获取。
AI 中文摘要
大型语言模型在通用代码生成方面表现出色,但处理科学代码的能力仍有待研究。现有数据集和基准测试在规模、领域覆盖或可执行验证方面存在局限性。为解决这些问题,我们提出了SciCodePile,这是迄今为止最大的科学代码语料库,由37737个公共存储库构建而成,涵盖128GB的代码,跨越多个计算科学学科。我们还精心策划了一个包含200个任务的可执行基准测试,每个任务都配备了沙盒执行环境和自动测试工具进行功能验证。我们在三个任务上评估了15个开源和闭源的大型语言模型,结果表明科学代码生成仍然极具挑战性。此外,我们还展示了SciCodePile在训练方面的效用,在我们的语料库上继续预训练可将科学代码完成的CodeBLEU提高2.84倍,在可执行基准测试上进行指令调整可将Pass@1提高4.79倍。所有代码和数据可通过给定的https链接获取。
英文摘要
Large language models (LLMs) excel at general-purpose code generation, yet how well they handle scientific code remains an open question. Existing datasets and benchmarks are limited in scale, domain coverage, or executable verification, leaving the true gap between current LLMs and reliable scientific code generators inadequately assessed. To address these limitations, we present SciCodePile, the largest scientific code corpus to date, constructed from 37,737 public repositories and collectively comprising 128GB of code that spans multiple computational science disciplines. From this corpus, we further curate an executable benchmark of 200 tasks, each equipped with a sandboxed execution environment and an automated test harness for functional verification. We evaluate 15 LLMs from both open-source and closed-source families on three tasks: prefix-to-suffix completion, fill-in-the-middle infilling, and executable code generation. Results show that scientific code generation remains highly challenging: The best CodeBLEU reaches only 38.13 and 38.37 on the two completion tasks, while the strongest model achieves just 12.30\% Pass@1 on the executable benchmark, underscoring how far current models remain from reliable scientific code generation. To demonstrate the training utility of SciCodePile, we further show that continued pretraining on our corpus improves CodeBLEU by $\times$2.84 on scientific code completion, and instruction tuning on our data improves Pass@1 by $\times$4.79 on the executable benchmark. All code and data are available at https://huggingface.co/SciCodePile.