arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.03592cs.SE

基于大语言模型的代码转换规则合成:潜力与局限

Code Transformation Rule Synthesis using LLMs: Potential and Limits

Axel Allain, Aymeric Blot, Djamel Eddine Khelladi, Mathieu Acher

首次发表
浏览论文内容

中文总结 AI 辅助

本研究评估GPT-5.4等三种LLMs在Comby等三种代码转换规则语言及六项软件演化任务中的表现,发现前沿LLMs合成转换规则潜力巨大,在正确性上优于反归约算法,仅适用性稍逊。

中文摘要 AI 辅助

由于大语言模型(LLMs)的黑箱特性,其可解释性有限且缺乏确定性,使用成本也会上升,尤其是在大型代码库的重复性任务中。为缓解这一问题,我们针对三种特定领域的转换规则语言(Comby、GritQL和Ast-Grep)开展了一项新颖的实证研究,在涵盖四项软件演化任务(API误用修正、程序修复、API迁移和语言版本迁移)的六个不同数据集上评估了三种LLMs(GPT-5.4、GPT-oss-120B和Llama3.1-8B)。结果表明,前沿模型的转换规则合成已超越概念验证阶段,GPT-5.4在大多数基准测试中始终保持较高的规则适用性,生成的转换最接近真实值;较小的开源模型GPT-oss-120B和Llama3.1-8B对更简单的局部变更仍有效,但在复杂迁移场景中表现吃力。我们还观察到,通过元变量的使用以及在许多数据集第一四分位数中较高的复用得分,模型具备不可忽视的泛化能力。最后,与反归约算法相比,LLMs在正确性上更优,但在规则适用性上表现更差。总体而言,我们的结果显示LLMs在生成合理、正确、可泛化且可复用的规则方面具有巨大潜力。

英文摘要

Due to their black-box nature, LLMs suffer from limited explain- ability and a lack of determinism. Their usage cost can also rise, particularly with repetitive tasks on large codebases. To mitigate this, we conduct a novel empirical study targeting three domain- specific languages for transformation rules, namely Comby, GritQL, and Ast-Grep. We evaluate three LLMs (GPT-5.4, GPT-oss-120B, and Llama3.1-8B) on six diverse datasets covering four software- evolution tasks: API misuse correction, program repair, API migra- tion, and language version migration. Our results provide evidence that transformation rule synthesis moves beyond proof-of-concept with strong frontier models. GPT-5.4 achieves consistently high rule applicability rates and produces transformations closest to the ground truth across most benchmarks. Smaller and open-weight GPT-oss-120B and Llama3.1-8B models remain effective for simpler, localized changes but struggle with complex migration scenarios. We also observe non-negligible generalizability through the usage of meta-variables and through a high reuse score in the first quartile of many datasets. Finally, when compared to the anti-unification algorithm, LLMs outperform it in correctness, but underperform in rule applicability. Overall, our results show great potential for LLMs to generate sound, correct, generalizable, and reusable rules.

发表机构

  • Univ Rennes, INRIA, CNRS, IRISA(雷恩大学、法国国家信息与自动化研究所、法国国家科学研究中心、IRISA)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑