WideSWE:编码智能体能否跨仓库协调更改?
WideSWE: Can Coding Agents Coordinate Changes Across Repositories?
- Zhejiang University(浙江大学)
- Tsinghua University(清华大学)
机构由 AI 辅助整理,请以论文原文为准。
中文总结 AI 辅助
提出WideSWE基准,评估编码智能体跨仓库协调更改的能力,包含120个真实任务。测试显示智能体成功率仅10.83%至42.50%,联合执行比独立执行更能利用相关仓库信息指导实现与验证。
中文摘要 AI 辅助
编码智能体的评估已从解决单个问题发展到执行长期开发任务,但任务完成情况仍主要是在单个代码库内进行评估。在软件生态系统中,许多功能和缺陷修复需要跨多个仓库进行协调更改。我们引入WideSWE来评估编码智能体在此类跨仓库任务上的表现。通过挖掘和审查103个软件生态系统中的更改,我们获得了120个真实世界任务,其中60个缺陷修复和60个功能任务保持平衡。我们从相关问题和拉取请求中生成提示。我们系统地审查并调整隐藏测试,以支持多样化的正确实现,同时保留所需行为和回归检查。在七种智能体配置中,完整任务成功率从10.83%到42.50%不等,其中将Codex CLI与GPT-5.6-sol配对的配置达到了最高成功率。轨迹显示,智能体未能识别必要的更改,识别了更改但未完成,或修改了所需仓库但未完全满足请求。为了检查一次处理一个仓库是否能缓解这些困难,我们将其与在相同提示下的联合执行进行了比较。独立执行主要恢复遗漏的工作,但在纠正先前尝试但未成功的实现方面效果较差。联合执行可以利用相关仓库的信息来指导实现和验证。代码可在以下网址获取:此https URL。
英文摘要
Coding-agent evaluation has progressed from resolving individual issues to carrying out long-horizon development, yet task completion is still largely assessed within a single codebase. In software ecosystems, many features and bug fixes require coordinated changes across multiple repositories. We introduce WideSWE to evaluate coding agents on such cross-repository tasks. Mining and reviewing changes across 103 software ecosystems yields 120 real-world tasks, balanced between 60 bug fixes and 60 features. We derive prompts from related issues and pull requests. We systematically review and adapt hidden tests to support diverse correct implementations while preserving required behavior and regression checks. Across seven agent configurations, full task success ranges from 10.83% to 42.50%, with the configuration pairing Codex CLI with GPT-5.6-sol achieving the highest rate. Trajectories show agents failing to identify necessary changes, recognizing changes but leaving them unfinished, or modifying the required repositories without fully satisfying the request. To examine whether working on one repository at a time can alleviate these difficulties, we compare it with joint execution under identical prompts. Independent execution mainly recovers omitted work and is less effective at correcting previously attempted but unsuccessful implementations. Joint execution can use information from related repositories to guide implementation and verification. Code is available at https://github.com/ZJU-ACES-ISE/WideSWE.