arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

性能优化基准测试是否可靠地衡量编码智能体?

Are Performance-Optimization Benchmarks Reliably Measuring Coding Agents?

Zhi Chen, Zhensu Sun, Yuling Shi, David Lo, Lingxiao Jiang

arXiv 2607.01211首次发表:更新:

发表机构

Singapore Management University; Shanghai Jiao Tong University(新加坡管理大学; 上海交通大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究审计了三个仓库级性能优化基准(GSO、SWE-Perf、SWE-fficiency),发现参考补丁的可复现性差、评分规则导致排名不一致,且多数任务已被公开提交解决,揭示了聚合排名掩盖的性能差距。

AI 中文摘要

仓库级性能优化基准测试(如GSO、SWE-Perf和SWE-fficiency)通过将补丁应用于真实仓库并与未优化基线和官方参考补丁比较运行时间来评估编码智能体。它们的排行榜分数越来越多地被用作编码智能体进展的证据,但这些分数可能混淆了运行时的不稳定性、基准特定的评分规则以及有多少任务已被至少一个公开提交解决。我们审计了这三个基准中的这些问题。首先,我们在四种常见的Google Cloud机器上重放了740个代码优化任务的官方参考补丁。大多数基准任务可以重放,但它们的参考补丁在每次跨机器重放中满足原始基准有效性规则的任务数量仅为:GSO 39/102,SWE-Perf 11/140,SWE-fficiency 411/498;SWE-Perf尤其脆弱,因为许多参考补丁产生接近零的运行时变化。其次,我们表明公开提交的排名强烈依赖于基准评分规则。在GSO和SWE-fficiency共享的八个公开提交中,官方排名在28对提交比较中有9对不一致,并且SWE-fficiency的排行榜评分规则为最差的十个任务分配了过高的分数权重(58.5%-82.8%)。第三,观察每个任务的10个公开提交,我们发现至少有一个提交在85.3%(384/450)的可重放GSO和SWE-fficiency任务中匹配或击败了参考补丁,并在99.8%(449/450)的任务中击败了未优化的基础代码。我们的研究通过识别具有更可靠性能信号的任务、量化每个任务的分数贡献以及揭示被聚合排名隐藏的剩余性能差距,补充了排行榜分数。

英文摘要

Repository-level performance-optimization benchmarks such as GSO, SWE-Perf and SWE-fficiency evaluate coding agents by applying patches to real repositories and comparing runtime against unoptimized baselines and official reference patches. Their leaderboard scores are increasingly used as evidence of coding-agent progress, but those scores can conflate runtime instability, benchmark-specific scoring rules, and how many tasks are already solved by at least one public submission. We audit these issues across the three benchmarks. First, we replay the official reference patches for 740 code optimization tasks across four common types of Google Cloud machines. Most benchmark tasks can be replayed, but their reference patches satisfy the original benchmark validity rules in every cross-machine replay for only 39/102 GSO tasks, 11/140 SWE-Perf tasks, and 411/498 SWE-fficiency tasks; SWE-Perf is especially fragile because many reference patches produce close-to-zero runtime changes. Second, we show that public submission rankings depend strongly on the benchmark scoring rule. Among eight public submissions shared by GSO and SWE-fficiency, the official rankings disagree on 9 of 28 pairwise submission comparisons, and SWE-fficiency's leaderboard scoring rule assigns the worst ten tasks overly high score weights of 58.5%-82.8%. Third, looking across 10 public submissions for each task, we find that at least one submission matches or beats the reference patch on 85.3% (384/450) of replay-valid GSO and SWE-fficiency tasks, and beats the unoptimized base code on 99.8% (449/450). Our study complements leaderboard scores by identifying tasks with more reliable performance signals, quantifying per-task score contributions, and exposing the remaining performance gaps that are hidden by aggregate rankings.

Comments12 pages, 7 figures. Public data: https://github.com/chenzhi-cz/performance-optimization-benchmark-reliability

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑