Zero2Repo:编码智能体能从零构建代码仓库吗?
Zero2Repo: Can Coding Agents Build Repositories from Scratch?
浏览论文内容
中文总结 AI 辅助
本文提出Zero2Repo基准,用于评估编码智能体从零构建代码仓库的能力,通过语言无关流水线生成任务并验证,发现最强智能体在11个任务中仅解决10个,失败多源于规范细节遗漏。
中文摘要 AI 辅助
编码智能体正越来越多地被要求构建软件而非修补软件,然而,用于从零构建代码仓库的基准测试大多局限于单一语言,并依赖于人工策划的任务。我们引入了Zero2Repo,这是一个基准测试,其中智能体接收产品需求文档、接口契约和空工作区,并且必须在项目的原生生态系统中交付一个完整的代码仓库。任务由一个与语言无关的创作流水线生成,该流水线将真实的、版本固定的开源项目转换为行为规范、可复现的环境和隐藏的验收测试。每个任务都通过执行来验证:源自上游项目的参考实现必须通过测试,并且对抗性验证必须表明测试会拒绝不正确的实现。评估在隔离容器中运行生产级编码智能体,在显式提交之前保留验收测试,并且仅当所有测试通过时才分配二元奖励,不涉及LLM评判。该流水线和测试框架不做出语言特定的假设,适用于主流编程生态系统;当前版本包含Python、TypeScript、Go和C++任务。即使在从前沿模型在训练期间极有可能见过的代码仓库中抽取的11个任务上,最强的智能体也仅解决了10个,并且每个失败的提交都通过了90-99%的隐藏测试;对于两个最强的智能体,67-100%的失败测试可追溯到规范中所述的单一遗漏或低频规则,而非缺失的子系统,因此每次失败都是一个具体的改进目标。
英文摘要
Coding agents are increasingly asked to build software rather than patch it, yet benchmarks for from-scratch repository construction are mostly limited to a single language and depend on manually curated tasks. We introduce Zero2Repo, a benchmark in which an agent receives a product requirements document, an interface contract, and an empty workspace, and must deliver a complete repository in the project's native ecosystem. Tasks are produced by a language-agnostic authoring pipeline that converts real, version-pinned open-source projects into behavioral specifications, reproducible environments, and hidden acceptance tests. Each task is validated by execution: a reference implementation derived from the upstream project must pass, and adversarial validation must show that the tests reject incorrect implementations. Evaluation runs production coding agents in isolated containers, withholds the acceptance tests until an explicit submission, and assigns a binary reward only when every test passes, with no LLM judge. The pipeline and harness make no language-specific assumptions and apply to mainstream programming ecosystems; the current release contains Python, TypeScript, Go, and C++ tasks. Even on 11 tasks drawn from repositories that frontier models have very likely seen during training, the strongest agent solves only 10, and every failing submission passes 90-99% of the hidden tests; for the two strongest agents, 67-100% of failed tests trace to a single omission or a low-frequency rule stated in the specification rather than to a missing subsystem, so each failure is a concrete target for improvement.
发表机构
- Gradient Data
- McGill University(麦吉尔大学)
- Carnegie Mellon University(卡内基梅隆大学)
- University of The Cumberlands(坎伯兰大学)
- Queen’s University(女王大学)
- MIT(麻省理工学院)
- University College London, University of London(伦敦大学学院)
- University of California, San Diego(加州大学圣地亚哥分校)
机构由 AI 辅助整理,请以论文原文为准。