arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SWE-NFI:面向非功能改进的编码智能体研究与基准测试

SWE-NFI: Studying and Benchmarking Coding Agents for Non-Functional Improvements

Pengyu Xue, He Yang Yuan, Xin Wang, Junkai Chen, Haonan Zhang, Boyuan Chen, Zishuo Ding, Zhenhao Li, Weiyi Shang

arXiv 2607.27409首次发表:更新:

发表机构

York University; Singapore Management University; University of Waterloo(约克大学; 新加坡管理大学; 滑铁卢大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出SWE-NFI基准,基于开源Python项目的真实合并拉取请求构建188个任务,评估编码智能体的非功能改进能力,发现智能体在该能力上普遍落后于人类开发者。

AI 中文摘要

尽管编码智能体在面向正确性的基准测试中取得了令人瞩目的性能,但它们在保持行为不变的前提下进行非功能改进(Non-Functional Improvements, NFIs)的能力仍未得到充分探索。在实际软件开发中,开发者会在不改变可观测行为的前提下持续提升软件质量,然而现有基准测试主要评估功能正确性,对这类非功能改进的评估支持有限。本文提出SWE-NFI,一个用于评估编码智能体在功能正确性之外的非功能改进能力的基准。该基准包含188个任务,这些任务源自开源Python项目中实际合并的拉取请求。我们将面向开发者的非功能改进转化为92条可执行规则,并开发了一套综合评估套件,结合功能正确性测试与基于规则的非功能改进评估。我们评估了当前最先进的商业及开源编码智能体,尽管表现最佳的智能体达到了70.0%的功能正确率,但所有被评估智能体在整体非功能改进能力上均普遍落后于人类开发者。在结构性代码改进方面,差距尤为明显:智能体的非功能改进得分范围为0.0至1.3,而人类基准得分为1.5。我们的基准与研究发现为评估和推进超越功能正确性的编码智能体提供了可复现的基础。

英文摘要

Although coding agents have achieved impressive performance on correctness-oriented benchmarks, their ability to make behavior-preserving non-functional improvements (NFIs) remains underexplored. In real-world software development, developers continuously improve software quality without changing observable behavior, yet existing benchmarks primarily evaluate functional correctness and provide limited support for assessing these non-functional improvements. In this paper, we present SWE-NFI, a benchmark for evaluating coding agents on NFIs beyond functional correctness. SWE-NFI contains 188 tasks constructed from real merged pull requests in open-source Python projects. We operationalize five developer-oriented NFI aspects into 92 executable rules and develop an evaluation suite that first applies task-specific functional preservation checks to assess whether key properties of the original code are preserved and then performs rule-based NFI evaluation. For outputs that pass these checks, we measure NFI improvement as the change in the normalized rule score for the task's target aspect relative to the original code. We evaluate eight commercial and open-source agent configurations and run each configuration five times on every task. At most 70.0% of valid agent outputs satisfy all task-specific functional preservation checks. Among outputs that pass these checks, NFI performance varies substantially across aspects. In particular, all evaluated configurations achieve mean Logic Patterns improvements between 0.0 and 1.3, below the Human Reference mean of 1.5. Our benchmark and findings provide a reproducible foundation for evaluating coding agents on non-functional improvements and identifying the NFI aspects where further advances are needed.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑