arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.11042cs.SEcs.AI

BackendForge:使用后端服务对智能端到端代码生成进行基准测试

BackendForge: Benchmarking Agentic End-to-End Code Generation with Backend Services

Yuzhe Guo, Mengzhou Wu, Yuan Cao, Jialei Wei, Dezhi Ran, Wei Yang, Tao Xie

首次发表
浏览论文内容

中文总结 AI 辅助

研究智能端到端代码生成评估问题,引入BackendForge基准,给定规范和契约让大语言模型生成容器化服务,用测试和代码代理共同进化测试预言机和参考服务,发现当前大语言模型虽能实现部分API行为,但生成完整后端服务仍有困难。

中文摘要 AI 辅助

大语言模型越来越多地用于智能编码环境,在此环境中它们可以检查文件、执行命令、运行测试、观察失败并迭代修改代码。这种转变引发了一个核心评估问题:智能大语言模型能否生成一个在执行时既可部署又行为正确的端到端软件工件?后端服务为该评估提供了一个可控但现实的基础。其API公开应用级可执行语义,并且可以通过黑盒HTTP交互根据OpenAPI契约确定性地检查部署行为。我们引入了BackendForge,这是一个由56个从真实开源应用程序改写的契约定义的后端生成任务的基准。给定一个可见规范和一个OpenAPI契约,大语言模型必须生成一个仅通过HTTP测试构建、部署和评估的容器化服务。为了在不引入隐藏要求的情况下加强评估,BackendForge使用测试代理和代码代理共同进化测试预言机和参考服务,其中测试代理提出基于规范的后端测试,代码代理修复参考实现。尽管性能最佳的模型GPT-5.5在基本预言机下55.4%的任务上取得成功,但在最终预言机下仅28.6%的任务上取得成功。这一差距表明当前的大语言模型可以实现许多本地API行为,但仍难以生成完整的后端服务。

英文摘要

Large language models (LLMs) are increasingly used in agentic coding settings, where they can inspect files, execute commands, run tests, observe failures, and iteratively revise code. This shift raises a central evaluation question: can an agentic LLM generate an end-to-end software artifact that is both deployable and behaviorally correct under execution? Backend services provide a controlled but realistic substrate for this evaluation. Their APIs expose application-level executable semantics, and deployed behavior can be checked deterministically against an OpenAPI contract through black-box HTTP interactions. We introduce BackendForge, a benchmark of 56 contract-defined backend generation tasks rewritten from real open-source applications. Given a visible specification and an OpenAPI contract, an LLM must generate a Dockerized service that is built, deployed, and evaluated only through HTTP tests. To strengthen evaluation without introducing hidden requirements, BackendForge uses a test agent and a code agent to co-evolve the test oracle and reference service, where the test agent proposes specification-grounded backend tests and the code agent repairs the reference implementation. Although the best-performing model, GPT-5.5, succeeds on 55.4\% of tasks under the base oracle, it succeeds on only 28.6\% under the final oracle. This gap suggests that current LLMs can implement many local API behaviors, but still struggle to produce complete backend services.

发表机构

  • Key Lab of HCST (PKU), MOE(北京大学计算机科学与技术国家重点实验室;软件学院,北京大学)
  • SCS, Peking University(北京通明湖信息技术应用创新中心)
  • Beijing Tongming Lake Information Technology Application Innovation Center (TLAIC)(复旦大学高级计算系统研究所)
  • Fudan University Institute of Systems for Advanced Computing(上海开放计算系统研究所)
  • Shanghai Institute of Systems for Open Computing

机构由 AI 辅助整理,请以论文原文为准。

↑