arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.30497cs.SE

Bridge:自动挖掘生态系统级API更新映射与客户端更新实例

Bridge: Automatically Mining Ecosystem-Scale API Update Mappings and Client Update Instances

Kai Gao, Yu Sun, Chang-ai Sun

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出客户端驱动框架Bridge,自动构建生态系统级库更新数据集,挖掘Java、Python的API更新映射与客户端更新实例,实验显示大语言模型在长尾映射的API推荐上表现较差。

中文摘要 AI 辅助

库更新通常需要调整客户端代码以适配API变更。API更新映射(用于标识遗留API与替代API之间的关系)、这些映射适用的版本转换、以及捕获具体API调用变更的客户端更新实例,是开发和评估自动化库更新技术的关键。现有的库演化数据集仅能捕获部分此类信息,且通常覆盖的第三方库数量较少。本文提出Bridge,这是一个客户端驱动的框架,用于自动构建连接API更新映射、版本转换和客户端更新实例的生态系统级库更新数据集。Bridge首先从客户端依赖更新提交中大规模挖掘候选更新实例,利用库侧证据对其进行验证,随后从已验证的实例中推导API更新映射,该设计确保每个保留的映射都至少基于一个客户端更新实例。在人工标注的基准数据集上,Bridge对Java的准确率达91.6%、召回率达88.7%,对Python的准确率达90.1%、召回率达64.0%。将其应用于WoC V3时,Bridge挖掘出381661个Java客户端更新实例和277259个Python客户端更新实例,分别对应2557个库中的18900个API更新映射,以及999个库中的4456个API更新映射。挖掘出的映射呈现明显的长尾分布,大多数映射仅出现在少数客户端更新实例中。作为该数据集的一项应用,我们评估了四个大语言模型在替代API推荐(库更新的关键步骤)上的表现,最佳推荐准确率对Java仅达37.1%,对Python仅达44.4%,且所有评估模型在频繁出现的映射上的表现,远优于仅在少数客户端更新实例中出现的映射,凸显了当前大语言模型在为长尾映射推荐替代方案时面临的困难。

英文摘要

Library updates often require adapting client code to API changes. API update mappings that identify relations between legacy and replacement APIs, version transitions that these mappings apply, and client update instances that capture concrete API call changes are essential for developing and evaluating automated library update techniques. Existing library evolution datasets capture only subsets of this information and typically cover few third-party libraries. In this paper, we present Bridge, a client-driven framework for automatically constructing ecosystem-scale library update datasets that connect API update mappings, version transitions, and client update instances. Bridge first mines candidate update instances from client dependency update commits at scale, validates them using library-side evidence, and then derives API update mappings from validated instances. This design grounds each retained mapping in at least one client update instance. On a manually annotated ground truth dataset, Bridge achieves 91.6% precision and 88.7% recall for Java and 90.1% precision and 64.0% recall for Python. Applied to WoC V3, Bridge mines 381,661 Java and 277,259 Python client update instances, representing 18,900 and 4,456 API update mappings across 2,557 and 999 libraries, respectively. The mined mappings exhibit a pronounced long-tail distribution, with most appearing in only a few client update instances. As one application of the dataset, we evaluate four large language models on replacement API recommendation, a key step in library updates. The best recommendation accuracy reaches only 37.1% for Java and 44.4% for Python, and all evaluated models perform substantially better on frequently observed mappings than on mappings observed in only a few client update instances, highlighting the difficulty current LLMs face in recommending replacements for mappings in the long tail.

发表机构

  • University of Science and Technology Beijing(北京科技大学)

机构由 AI 辅助整理,请以论文原文为准。

↑