arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Prefactory:库采用重构的自动发现与应用

Prefactory: Automated Discovery and Application of Library-Adoption Refactorings

Islem Bouzenia, Michael Pradel

arXiv 2607.17211首次发表:更新:

AI 中文总结

研究如何自动进行库采用重构。核心方法是用 LLM 合成搜索启发式方法,收集元数据和词汇生成检测器找候选函数,排序后用 LLM 生成重构并验证。主要贡献是在 PrefactoryBench 上表现优于基线,产生多个经测试验证的重构。

AI 中文摘要

用库 API 调用替换手写代码是常见的重构方式,可减少代码量、使代码更地道并重用经过充分测试的实现。然而,许多库采用机会难以自动发现,现有工具覆盖模式有限,基于大语言模型(LLM)的方法成本高、难重现和规模化应用。本文介绍了 Prefactory,一种用于 Python 中库采用重构的自动化方法。其关键思想是使用 LLM 合成可执行的搜索启发式方法,而非依赖对代码库的重复 LLM 提示。给定目标项目和库名称,Prefactory 收集库元数据和项目词汇,生成检测器,在扫描阶段找到候选函数,对候选函数进行启发式排序,用 LLM 为排名最高的函数生成重构,并通过项目测试和新生成的差异测试验证结果。我们在 PrefactoryBench 上评估了 Prefactory,它在文件级别检测到 75 个实例,函数级别检测到 56 个实例,而最强基线(Codex CLI)分别为 35 个和 32 个。从 56 个检测到的函数中,Prefactory 产生了 40 个经测试验证的重构。

英文摘要

Replacing hand-written code with library API calls is a common refactoring that can reduce code size, make code more idiomatic, and reuse well-tested implementations. Yet many library-adoption opportunities are hard to find automatically: the original code often does not mention the target library and may resemble the library API only in behavior, with little syntactic overlap. Existing tools, such as linters and static modernizers, cover only a small set of manually specified patterns. LLMs and LLM-based agents, on the other hand, can generalize to more patterns, but they are costly, difficult to reproduce and to apply systematically at scale. This paper introduces Prefactory, an automated approach for library-adoption refactoring in Python. The key idea is to use an LLM to synthesize executable search heuristics rather than relying on repeated LLM prompting over a codebase. Given a target project and a target library name, Prefactory collects library metadata and project vocabulary, then generates lexical and structural detectors. Prefactory executes the detectors during a scan phase to find candidate functions. It then heuristically ranks the candidate functions, generates refactorings for the highest-ranked ones using an LLM, and validates the results with project tests and newly generated differential tests. We evaluate Prefactory on PrefactoryBench, a benchmark of 100 real-world library-adoption refactorings from 61 open-source Python projects and 18 libraries. Prefactory detects 75 instances at the file level and 56 at the function level, compared with 35 and 32 for the strongest baseline (Codex CLI). From the 56 detected functions, Prefactory produces 40 test-validated refactorings.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑