arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.04167cs.SEcs.AI

SWE-Gate:软件工程智能体仅通过功能测试是不够的

SWE-Gate: Passing Functional Tests Is Not Enough for Software Engineering Agents

  • Sun Yat-sen University(中山大学)
  • Zhejiang University(浙江大学)
  • Chongqing University(重庆大学)

机构由 AI 辅助整理,请以论文原文为准。

Xin He, Yanlin Wang, Mingwei Liu, Jiachi Chen, Hongyu Zhang, Guanbin Li

中文总结 AI 辅助

SWE-Gate基准结合功能测试与审查约束评估,发现仅通过功能测试的软件工程智能体中约34%未满足审查约束,仅功能评估会高估智能体的修复能力。

中文摘要 AI 辅助

仓库级软件工程基准已显著推进了编码智能体的评估,但现有基准主要衡量生成的补丁是否通过功能测试,却忽略了实际软件开发中常影响补丁是否被接受的、源自代码审查的接受约束(审查约束)。我们提出SWE-Gate,这是一款用于软件工程智能体的仓库级基准,它在评估功能正确性的同时,明确评估对审查约束的遵守情况。SWE-Gate从真实的拉取请求审查评论中提取审查约束,并围绕这些约束合成仓库级修复实例。每个实例都提供独立的功能测试和约束测试,以及不合规补丁和黄金补丁,从而能明确区分问题解决能力与审查约束遵守能力。我们构建的SWE-Gate包含303个仓库级修复实例,涵盖75个不同软件领域的开源Python仓库。在通用编码智能体框架下,对四种不同能力层级的大语言模型后端进行的实验显示,功能成功与完整修复规范下的成功之间存在显著差距:在644个通过功能测试的修复中,有221个未能满足提供的审查约束。这些发现表明,仅基于功能的评估会高估智能体满足仓库级修复任务全部要求的能力。包含代码、数据和实验结果的复现包可在该https URL获取。

英文摘要

Repository-level software engineering benchmarks have significantly advanced the evaluation of coding agents, but existing benchmarks primarily measure whether generated patches pass functional tests and overlook review-derived acceptance constraints (review constraints) that often influence whether a patch is acceptable in real-world software development. We introduce SWE-Gate, a repository-level benchmark for software engineering agents that explicitly evaluates review constraint compliance alongside functional correctness. SWE-Gate derives review constraints from real pull request review comments and synthesizes repository-level repair instances around these constraints. Each instance provides separate functional and constraint tests, together with non-compliant and gold patches, enabling explicit separation between issue resolution capability and review constraint compliance. We construct SWE-Gate with 303 repository-level repair instances spanning 75 open-source Python repositories across diverse software domains. Experiments with four LLM backends spanning different capability levels under a common coding-agent scaffold reveal a substantial gap between functional success and success under the complete repair specification: among 644 repairs that pass the functional tests, 221 fail to satisfy the provided review constraints. These findings show that functional-only evaluation overestimates agents' ability to satisfy the full requirements of repository-level repair tasks. The replication package including code, data, and experimental results is available at https://github.com/DeepSoftwareAnalytics/SWE-Gate.

补充信息

↑