arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

E2E-SWE:基准测试大语言模型从零构建可用代码库

E2E-SWE: Benchmarking LLMs on Building Working Codebases from Scratch

Hantian Ding, Chloe Bi, Jiacheng Zhu, John Yang, Matt Deitke, Pengcheng Yin, Zijian Wang, Rui Hou

arXiv 2609.38335首次发表:更新:

发表机构

Meta Superintelligence Labs(元超级智能实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该基准E2E-SWE通过186个跨11种语言的整仓库生成任务评估LLM智能体从零构建可用代码库的能力,发现pass@1在11.7%至67.7%间波动,凸显系统级推理需求。

AI 中文摘要

由大语言模型(LLM)驱动的编码智能体正从进行局部代码修改演进到开发完整的软件仓库。然而,评估仓库规模的代码生成仍具挑战性:任务必须要求系统级推理,同时确保所有被评估的行为都被精确指定且不依赖于任何特定实现。我们引入了E2E-SWE,一个用于评估编码智能体能否端到端构建完整、功能可用的软件仓库的基准。E2E-SWE包含186个跨11种编程语言的整仓库生成任务。给定仅一份自然语言规格说明和一个空工作区,智能体必须实现一个完整的、可安装的项目,以满足一套全面的隐藏测试。每个任务由软件工程师与LLM协作构建;他们共同开发测试套件和相应的与实现无关的规格说明。为确保任务被良好指定且实际可解,我们进一步对任务进行迭代验证过程,其中自主智能体使用静态检查和从真实模型运行中观察到的失败来审计和修复任务缺陷。评估13个前沿模型,我们发现端到端仓库生成能力存在显著差异,pass@1从11.7%到67.7%不等,提供了强大的模型区分度,同时为未来进展留下了充足空间。对智能体轨迹的分析进一步揭示了长且前置的推理模式,突出了从零构建可用代码库所需的规划和系统级推理。

英文摘要

Coding agents powered by large language models (LLMs) are evolving from making localized code changes to developing complete software repositories. However, evaluating repository-scale generation remains challenging: tasks must demand system-level reasoning while ensuring that all evaluated behaviors are precisely specified and independent of any particular implementation. We introduce E2E-SWE, a benchmark for evaluating whether coding agents can build complete, functional software repositories end to end. E2E-SWE contains 186 whole-repository generation tasks spanning 11 programming languages. Given only a natural-language specification and an empty workspace, an agent must implement a complete, installable project that satisfies a comprehensive suite of hidden tests. Each task is constructed by a software engineer in collaboration with an LLM; together, they develop the test suite and a corresponding implementation-independent specification. To ensure that tasks are well specified and practically solvable, we further subject them to an iterative verification process in which autonomous agents audit and repair task defects using static inspection and failures observed from real model rollouts. Evaluating 13 frontier models, we find substantial variation in end-to-end repository generation ability, with pass@1 ranging from 11.7% to 67.7%, providing strong model differentiation while leaving considerable headroom for future progress. Analysis of agent trajectories further reveals long, front-loaded reasoning patterns, highlighting the planning and system-level reasoning required to construct working codebases from scratch.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑