发表机构
NVIDIA; University College London; New York University; Georgia Institute of Technology; Cornell University; Stanford University; Northeastern University(英伟达; 伦敦大学学院; 纽约大学; 佐治亚理工学院; 康奈尔大学; 斯坦福大学; 东北大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对语言模型代理在多漏洞软件维护中的评估问题,引入ChainSWE基准,通过收集Python项目问题链进行评估,发现链长度增加时性能下降。
AI 中文摘要
语言模型代理越来越多地用于长期维护代码库,修复相关缺陷流并传递上下文。现有软件工程基准每次评估一个错误,忽略了累积依赖性。本文引入ChainSWE,首个用于评估共享代码库中顺序、相关错误修复的代理基准,评估发现链长度增加时性能下降高达70%。
英文摘要
Language model (LM) agents are increasingly deployed to maintain codebases over extended periods, fixing streams of related defects while carrying context from one fix to the next. Yet existing software engineering (SWE) benchmarks evaluate models one bug at a time: the repository is reset, the codebase is re-read, and a single self-contained issue is graded in isolation. This setting collapses a continuous maintenance workflow into a series of independent sessions, ignoring the cumulative dependencies that make real-world bug fixing challenging. To bridge this gap, we introduce ChainSWE, the first benchmark for evaluating agents on sequential, dependent bug fixes within a shared codebase. We collect chronological chains of 304 issues across 54 Python projects, mined from six SWE-bench-family datasets. Our evaluation across a range of agents and models reveals a consistent performance drop by up to 70% as the chain length increases.