arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SWE-Bench ProMax:面向大规模多语言代码重构的智能体基准测试

SWE-Bench ProMax: Benchmarking Agents on Large-Scale Multilingual Code Refactoring

Yuling Shi, Jinghan Xu, Kelin Fu, Wenhao Zeng, Shilin He, Lei Zhang, Yue Liu, Zelin Zhao, Terry Yue Zhuo, Jialun Cao, Siyu Ye, Tianyu Liu, Kai Cai, Shing-Chi Cheung, Xiaodong Gu

arXiv 2608.09802首次发表:更新:

发表机构

Shanghai Jiao Tong University; Peking University; The Hong Kong University of Science and Technology; Douyin Group; University of Chinese Academy of Sciences; National University of Singapore; Monash University(上海交通大学; 北京大学; 香港科技大学; 抖音集团; 中国科学院大学; 新加坡国立大学; 莫纳什大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究推出SWE-Bench ProMax多语言代码重构基准,含170个跨7种语言的大规模重构任务,经严格筛选后前沿模型仅达41.2%解决率,为AI编码智能体提供挑战性测试。

AI 中文摘要

随着AI编码智能体承担日益复杂、长周期的软件工程任务,现有基准测试已迅速饱和,其评估质量也受到严重质疑:一项近期审查发现,近60%未解决的SWE-bench Verified实例存在缺陷测试——要么是过于狭隘、会拒绝正确解决方案的测试,要么是过于宽泛、会检查未明确要求的测试;此外,前沿模型还能逐字重现训练数据中的最优补丁。代码重构需要跨多个文件进行协调一致、保持行为不变的变更,这对智能体能力而言是一个难度大得多、也更贴近现实的测试,但当前基准测试对该场景的覆盖仍不足。我们推出SWE-Bench ProMax,这是一个由专家精心筛选的多语言代码重构基准测试,包含170个实例,来自7种编程语言(Python、Java、TypeScript、Go、C、C++和Rust)的真实提交。每个实例都经过严格的多阶段筛选,直接解决了先前基准测试中发现的质量问题:问题描述被重新编写,以提供精确、无歧义的规范;测试套件经过人工审查,以删除过于狭隘和过于宽泛的测试。复杂度不足或跨文件范围有限的任务被过滤,最终得到的基准测试包含具有挑战性的大规模重构任务,每个实例平均修改11.4个文件、261.6行代码,规模远超现有基准测试。对前沿模型在两种智能体框架下的实验表明,表现最好的模型仅达到41.2%的解决率,证实SWE-Bench ProMax对当前AI编码智能体构成了有意义且未饱和的挑战。我们的基准测试可在该https URL获取。

英文摘要

As AI coding agents take on increasingly complex, long-horizon software engineering tasks, existing benchmarks are rapidly saturating and their evaluation quality has come under serious scrutiny: a recent audit found that nearly 60% of unsolved SWE-bench Verified instances contain flawed tests -- either overly narrow tests that reject correct solutions or overly broad tests that check unstated requirements -- and that frontier models can verbatim reproduce gold patches from training data. Code refactoring, which requires coordinated, behavior-preserving changes across many files, offers a substantially harder and more realistic test of agent capability, yet remains underserved by current benchmarks. We introduce SWE-Bench ProMax, an expert-curated, multilingual code refactoring benchmark of 170 instances drawn from real commits across seven programming languages (Python, Java, TypeScript, Go, C, C++, and Rust). Every instance undergoes rigorous, multi-stage curation that directly addresses the quality problems identified in prior benchmarks: issue descriptions are rewritten from scratch to provide precise, unambiguous specifications, and test suites are manually reviewed to remove overly narrow and overly broad tests. Tasks with insufficient complexity or limited cross-file scope are filtered out, yielding a benchmark of challenging, large-scale refactoring tasks that average 11.4 modified files and 261.6 lines of code per instance, substantially exceeding the scale of existing benchmarks. Experiments with frontier models under two agent scaffolds show that the best model achieves only 41.2% resolve rate, confirming that SWE-Bench ProMax presents a meaningful and unsaturated challenge for current AI coding agents. Our benchmark is available at https://huggingface.co/datasets/swe-bench-promax/SWE-Bench-ProMax.

CommentsPublished as a conference paper at COLM 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑