arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.17957cs.SE

DepRepair:基于大语言模型的针对依赖项破坏更改的源代码修复

DepRepair: LLM-Based Source-Code Repair for Dependency Breaking Changes

Shenghao Yang, Bo Lu, Yaochen Liu, Yu Kang, Qiongfang Zhang, Chetan Bansal, Saravan Rajmohan, Minghua Ma

AI总结:

研究软件项目因第三方库更新致依赖项破坏更改的问题,提出DepRepair单调用大语言模型方法,通过证据过滤器、用法定位器和子类别感知指南,基于结构化上游证据修复,在DepBench基准测试中取得高可执行通过率。

AI中文摘要:

现代软件项目依赖众多第三方库,其更新常带来破坏更改。使消费代码适应此类更改既费力又易出错。现有工作要么描述依赖项破坏更改却不生成经过验证的消费端补丁,要么仅在目标存储库包含失败和修复上下文的设置中研究自动修复。然而,依赖项破坏更改违反此假设,决定性修复证据在上游发布说明和API差异中,且无失败测试能定位消费者代码何处出错,导致修复信息不足。为在真实数据上研究此跨存储库问题,我们引入DepBench,一个包含四个生态系统中95个真实世界依赖项更新实例的基准,每个实例都配有基于Docker的可执行预言机来运行消费者自身测试。为应对这些挑战,我们提出DepRepair,一种单调用大语言模型方法,通过三个组件在结构化上游证据基础上进行修复:一个提炼相关上游更改的证据过滤器、一个识别受影响消费端位置的用法定位器以及一个根据破坏更改类型定制修复的子类别感知指南。在DepBench上评估,DepRepair在每个主干上获得最高可执行通过率,使用GPT-5.5时达到89.5%,使用Claude Opus 4.6时达到82.1%。我们还发现原始上游证据会使大语言模型和智能体通过率降低7 - 23个百分点,而结构化证据则持续提高它们。

英文摘要:

Modern software projects depend on numerous third-party libraries, whose updates often introduce breaking changes. Adapting consumer code to such changes remains labor-intensive and error-prone. Existing work either characterizes dependency breaking changes without producing a verified consumer-side patch, or studies automated repair only in settings where the failure and repair context are contained within the target repository. However, dependency breaking changes violate this assumption: the decisive repair evidence lies upstream in release notes and API diffs, and no failing test localizes where the consumer breaks, leaving the repair under-informed. To study this cross-repository problem on real data, we introduce DepBench, a benchmark of 95 real-world dependency-update instances across four ecosystems, each paired with a Docker-based executable oracle that runs the consumer's own tests. To address these challenges, we propose DepRepair, a single-call LLM approach that grounds repair in structured upstream evidence through three components: an evidence filter that distills relevant upstream changes, a usage locator that identifies affected consumer sites, and a subcategory-aware guide that tailors repairs to the breaking-change type. Evaluated on DepBench, DepRepair attains the highest executable pass rate on each backbone, achieving 89.5% with GPT-5.5 and 82.1% with Claude Opus 4.6. We further find that raw upstream evidence reduces LLM and agent pass rates by 7--23 percentage points, whereas structured evidence consistently improves them.

补充信息

↑