来自地狱的更新:编码智能体能否在依赖升级的隐藏损坏中存活?
Update from Hell: Can Coding Agents Survive Hidden Breakage in Dependency Upgrades?
- Monash University(莫纳什大学)
- Peking University(北京大学)
- Microsoft(微软)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文推出含203项真实依赖升级任务的DEPBENCH基准,评估主流编码智能体后发现其仅解决51.2%任务,凸显当前智能体能力与软件维护需求的差距。
AI中文摘要:
现代软件系统高度依赖第三方依赖库,但升级这些依赖库仍是一项成本高昂的维护活动。依赖升级并不总能保留现有代码所假定的函数签名、类型系统、应用程序编程接口(API)或运行时语义。因此,开发者常常需要对源代码进行适配,以应对依赖库引发的变更。然而,这类代码层面的变更往往未明确传达给项目维护者,对软件可靠性构成重大挑战。与此同时,编码智能体作为一种新型软件开发工具应运而生,凭借其自动化能力正被开发者日益采用。本文中,我们推出DEPBENCH,这是一个包含203项真实依赖升级任务的基准测试,涵盖五个包生态系统,涉及五个语言社区,每项任务都包含需要源代码适配的隐藏代码层面变更。我们在DEPBENCH上评估主流编码智能体,最佳完成配置仅解决了203项任务中的104项(51.2%),不同智能体框架、模型和生态系统间存在显著差异,凸显了当前智能体能力与现实世界软件维护需求之间的重要差距。
英文摘要:
Modern software systems rely heavily on third-party dependencies, but upgrading those dependencies remains a costly maintenance activity. Dependency upgrades do not always preserve the function signatures, type systems, APIs, or runtime semantics assumed by existing code. Consequently, developers often need to perform source code adaptations to accommodate dependency-induced changes. However, such code-level changes are often not explicitly communicated to project maintainers, posing a significant challenge to software reliability. Meanwhile, coding agents have emerged as a new form of software development tool and are increasingly adopted by developers due to their automation capabilities. In this paper, we introduce DEPBENCH, a benchmark consisting of 203 real-world dependency-upgrade tasks across five package ecosystems spanning five language communities, each involving hidden code-level changes that require source code adaptation. We evaluate mainstream coding agents on DEPBENCH. The best completed configuration solves only 104/203 tasks (51.2%), with substantial variation across agent harnesses, models, and ecosystems, highlighting an important gap between current agent capabilities and real-world software maintenance needs.