谁破坏了我?行为依赖断裂的执行引导修复
Who Broke Me? Execution-Guided Repair of Behavioral Dependency Breaks
- The University of Melbourne(墨尔本大学)
- Sun Yat-sen University(中山大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对依赖升级导致的行为性破坏难以定位根API的问题,提出BBCFixer方法,通过新旧版本下运行测试并排序差异调用识别根API,结合BBCBench基准验证,显著提升修复通过率并减少步骤。
AI中文摘要:
依赖升级可能在未改变库接口的情况下破坏下游项目。此类行为性破坏变更对开发者而言难以修复,因为失败的测试并不总是指向根API,即导致破坏的上游API。现有的基于LLM的修复方法从编译器反馈或库文档中获取修复证据。然而,行为性破坏不会产生编译器反馈,且通常没有文档记录。未收到升级证据的智能体通常也不会自行检索上游证据,其大多数失败修复未能识别根API。库差异提供了有用的上游证据,但必须先识别根API才能选择差异中的相关部分。我们提出BBCFixer,一种修复方法,它在旧版和新版库版本下运行失败测试,对返回值不同的调用进行排序以识别候选根API,并用该候选过滤库差异。为评估行为性破坏的修复,我们进一步引入BBCBench,一个包含100个Python和JavaScript行为性依赖破坏的基准。我们在BBCBench上将BBCFixer与两个基线进行比较:(1)Pure Agent,一个未接收升级证据的智能体,和(2)BDUpdater,一种挖掘库文档的方法。BBCFixer相对于Pure Agent将平均通过率提高了最多17%,且每次成功修复所需的智能体步骤最多减少27%。因此,BBCFixer提供了一种选择行为性破坏背后上游证据的方法,使LLM智能体能够更有效且高效地修复行为性破坏变更。
英文摘要:
Dependency upgrades can break downstream projects without changing the library interface. Such behavioral breaking changes are difficult for developers to fix, because the failing test does not always point to the root API, the upstream API that causes the break. Existing LLM-based repair methods obtain evidence for the repair from compiler feedback or library documentation. However, a behavioral break produces no compiler feedback and is often undocumented. Agents that receive no upgrade evidence also usually do not retrieve upstream evidence themselves, and most of their failed repairs do not identify the root API. The library diff provides useful upstream evidence, but the root API must first be identified to select the relevant part of the diff. We present BBCFixer, a repair method that runs the failing test under the old and new library versions, ranks the calls whose return value differs to identify a candidate root API, and filters the library diff with that candidate. To evaluate repairs of behavioral breaks, we further introduce BBCBench, a benchmark of 100 behavioral dependency breaks in Python and JavaScript. We compare BBCFixer on BBCBench with two baselines: (1) Pure Agent, an agent that receives no upgrade evidence, and (2) BDUpdater, a method that mines library documentation. BBCFixer raises the average pass rate by up to 17% relative to Pure Agent and needs up to 27% fewer agent steps per successful repair. BBCFixer thus provides a way to select the upstream evidence behind a behavioral break, so that an LLM agent can repair behavioral breaking changes more effectively and efficiently.