AI 中文总结
研究对比新旧模型在存储库级修复任务中生成补丁的非功能性质量,通过两个案例研究,用CodeQL等工具评估,发现新模型解决实例更多,但非功能性指标无一致改善,旨在推动对模型软件工程能力的全面评估。
AI 中文摘要
存储库级编码基准通常通过比较新旧模型的解决率来衡量模型能力的进展。然而,这种关注忽略了所生成补丁的非功能性质量在不同模型代际间是否也发生了变化。本研究调查了在可比的存储库级修复任务中,较新模型生成的功能正确补丁在非功能性特征方面是否优于较早模型。我们在SWE-bench Lite上对四个Claude和DeepSeek模型进行了两个案例研究。使用相同的SWE-agent功能修复设置,我们用CodeQL、CodeScene、CPU时间和峰值内存来评估生成的补丁。主要分析比较了常见解决实例上的模型。静态分析结果表明,大多数CodeQL配对差异为零,经过Holm校正后,没有CodeQL或CodeScene比较仍具有显著性。CPU时间差异较小且在模型家族间不一致,而在基准测试工作负载下,较新模型的峰值内存使用略高,绝对差异较小。各个CodeQL规则和CodeScene类别的差异在模型家族间各不相同,且未通过多重比较校正。总体而言,较新模型解决了更多实例,但在两个模型都解决的任务上,所测量的非功能性指标没有一致的改善。通过这项研究,我们希望鼓励对模型的实际软件工程能力进行更全面的评估。
英文摘要
Repository-level coding benchmarks typically measure progress in model capability by comparing the resolved rates of later and earlier models. However, this focus overlooks whether the non-functional quality of their generated patches has also changed across model generations. This study investigates whether later models produce functionally correct patches with better non-functional characteristics than earlier models on comparable repository-level repair tasks. We conducted two case studies involving four Claude and DeepSeek models on SWE-bench Lite. Using the same SWE-agent functional repair setting, we evaluated the generated patches with CodeQL, CodeScene, CPU time, and peak memory. Our primary analysis compared the models on commonly resolved instances. The static analysis results showed that most CodeQL paired differences were zero and that no CodeQL or CodeScene comparison remained significant after Holm correction. CPU time differences were small and inconsistent across model families, while peak memory usage was slightly higher for the later models under the benchmark test workload, with small absolute differences. Differences in individual CodeQL rules and CodeScene categories varied across model families and did not survive multiple-comparison correction. Overall, later models resolved more instances but showed no consistent improvement in the measured non-functional indicators on tasks solved by both models. Through this study, we hope to encourage a more comprehensive evaluation of models' practical software engineering capabilities.
CommentsAccepted in APSEC 2026 Techincal Track