发表机构
Imperial College London; The Hong Kong Polytechnic University; Korea University; Jinan University(帝国理工学院; 香港理工大学; 高丽大学; 暨南大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文审计SWE-bench排行榜,发现顶尖编码智能体成功高度共享且分数受模型-脚手架配对影响,导致微小差异无法支持可靠排序,并提出五步审计协议及报告比较集特定解决率以替代排名解读。
AI 中文摘要
编码智能体排行榜上的微小差异常被解读为系统间的排序。我们通过使用254个SWE-bench提交(涵盖四个分割)而不运行模型,来审计已发布的结论是否支持这种解读。在Verified分割上,领先的两个条目各自解决了500个实例中的396个。前十名共享285个成功和51个失败,留下164个实例来区分它们的结果。前沿解集的中位嵌套度为0.935,而分数隐含基线为0.774,表明成功高度共享。分数还取决于所评估的模型-脚手架配对:观察到的模型内脚手架范围达到29.8个百分点,而前三十名的差距为8.8个百分点。九项单元均值交互检验中有六项在霍尔姆校正后仍显著,尽管此观察性设计无法识别因果性脚手架效应。精确配对麦克尼马尔检验在alpha=0.05下未能区分Verified前三十名中29个相邻配对中的任何一个,而较大的Test分割则区分了23个中的14个。基于领先者的规则产生三个描述性层级,或经霍尔姆校正后为两个;未拒绝并不等同于等价。我们发布了分区和五步审计协议,该协议分析共享结果、检验配对差异、报告分组敏感性,并估计解决所需的实例预算。这些结果促使报告比较集特定的解决率和模型-脚手架来源,而非将小的总体差距解释为已确立的排名差异。
英文摘要
Small differences on coding-agent leaderboards are often read as an ordering of systems. We audit whether the published verdicts support this reading, using 254 SWE-bench submissions across four splits without running models. On Verified, the leading two entries each resolve 396 of 500 instances. The top ten share 285 successes and 51 failures, leaving 164 instances that distinguish their outcomes. Frontier solution sets have median nesting 0.935 against a score-implied baseline of 0.774, indicating strongly shared successes. Scores also depend on the evaluated model-scaffold pair: observed within-model scaffold ranges reach 29.8 percentage points, compared with the 8.8-point spread of the top thirty. Six of nine cell-mean interaction tests remain significant after Holm correction, although this observational design does not identify causal scaffold effects. Exact paired McNemar tests separate none of the 29 adjacent Verified top-thirty pairs at alpha=0.05, while the larger Test split separates 14 of 23. A stated leader-based rule yields three descriptive tiers, or two after Holm correction; non-rejection does not establish equivalence. We release the partition and a five-step audit protocol that profiles shared outcomes, tests paired differences, reports grouping sensitivity, and estimates the instance budget needed for resolution. The results motivate reporting comparison-set-specific resolution and model-scaffold provenance instead of interpreting small aggregate gaps as established rank differences.
CommentsAccepted at ADMA 2026 (International Conference on Advanced Data Mining and Applications), Special Session on Responsible Data Intelligence. Camera-ready version, 15 pages, 4 figures