arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.19799cs.CLcs.SE

SWE-bench Science:编码智能体能解决科学领域的工程任务吗?

SWE-bench Science: Can Coding Agents Resolve Engineering Tasks in Science?

Zhipeng Xu, Jiahao Lu, Yining Zheng, Yuxin Wang, Xipeng Qiu

首次发表
浏览论文内容

中文总结 AI 辅助

本研究推出SWE-bench Science基准,含20个领域98个仓库的119项任务,评估发现最优智能体pass@1低于50%,揭示科学软件工程的挑战,分析四类失败机制并经消融实验明确科学知识的作用,为编码智能体研究提供测试平台。

中文摘要 AI 辅助

软件日益成为科学仪器本身的组成部分,科学代码的故障不仅会影响程序行为,还可能损害科学结论背后的证据。然而,现有的编码智能体评估大多侧重于总体任务成功率,难以深入理解智能体在修复科学软件时失败的原因。我们推出了SWE-bench Science,这是一个针对科学软件工程的仓库级基准,包含来自20个科学领域的98个GitHub仓库的119个任务。每个任务被分为三类范式:问题驱动型、专家探索型和工程集成型。即使是表现最佳的智能体Claude Code with Opus-5(max),其pass@1也低于50%,凸显了科学软件工程带来的巨大挑战。我们在分析中确定了四种反复出现的失败机制:科学知识或抽象能力不足、探索方向错误或仅进行表面修复、修复覆盖不完整或系统集成不足、以及无法将科学知识推广到观察到的案例之外。我们还进行了配对消融实验,移除明确的科学指导,同时保留仓库和可执行的工程上下文。结果显示,科学知识并非始终有益:合理的信息可以限制修复,提高平均性能和令牌效率,而对齐不佳的指导会导致锚定效应,不一定能提高精确修复成功率。总之,SWE-bench Science为研究编码智能体在科学软件工程中的能力和失败机制提供了广泛的测试平台。

英文摘要

Software increasingly functions as part of the scientific instrument itself, making failures in scientific code capable of compromising not only program behavior but also the evidence underlying scientific conclusions. Yet existing evaluations of coding agents largely emphasize aggregate task success, providing limited insight into why agents fail when repairing scientific software. We introduce \textbf{SWE-bench Science}, a repository-level benchmark for scientific software engineering comprising 119 tasks from 98 GitHub repositories across 20 scientific domains. Each task is organized into one of three paradigms: Issue-driven, Expert-exploratory, and Engineering-integration. Even the best-performing agent, \textbf{Claude Code with Opus-5 (max), achieves a pass@1 below 50\%}, highlighting the substantial challenges posed by scientific software engineering. We identify four recurring failure mechanisms: deficits in scientific knowledge or abstraction, misguided exploration or surface-level repair, incomplete repair coverage or system integration, and failures to generalize scientific knowledge beyond observed cases in our analysis. We further conduct a paired ablation that removes explicit scientific guidance while preserving the repository and executable engineering context. The results show that scientific knowledge is not uniformly beneficial: well-grounded information can constrain repair and improve average performance and token efficiency, whereas poorly aligned guidance can induce anchoring and does not necessarily improve exact repair success. Together, SWE-bench Science provides a broad testbed for studying both the capabilities and failure mechanisms of coding agents in scientific software engineering.

补充信息

↑