发表机构
Sapienza Università di Roma; KU Leuven(罗马大学; 荷语鲁汶大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对大语言模型时代模型驱动工程(MDE)的价值问题,提出融合MDE可靠性与LLM易用性的多智能体框架ARTHUR,用于重构MDE生成的遗留Java代码,实验验证了MDE仍具重要价值。
AI 中文摘要
在软件开发与大语言模型(Large Language Models,LLM)深度绑定的时代,模型驱动工程(Model Driven Engineering,MDE)是否仍有意义?这引出了一个问题:MDE与LLM方法能够在多大程度上成功结合,以弥补各自方法的不足。本文通过研究一个核心问题来尝试回答该疑问:由大语言模型驱动的智能体能否改进模型驱动工程(MDE)工具从UML图和文本规格说明中生成的代码?为探究这一问题,我们开发了一种全新的由LLM驱动的多智能体方法,名为ARTHUR(基于混合UML推理的架构重构,Architecture Refactoring Through Hybrid UML Reasoning),该框架旨在结合MDE的可靠性与LLM的易用性。ARTHUR的设计目标是对传统基于规则的MDE代码生成器从UML图生成的遗留Java代码进行重构。为评估问题的答案,我们使用为此研究构建的自定义数据集,对多个项目的遗留代码进行了重构。ARTHUR能够添加对Spring Boot等现代框架的支持,同时确保符合基于模型的测试技术要求,以验证代码仍符合初始模型的规格说明。随后我们从时间与成本、测试通过率,以及compile@k、pass@k和pass^k指标方面对所得结果进行了衡量。我们还观察了不经过重构、直接从概念模型生成代码的效果。我们的初步测试结果表明,MDE远未到被淘汰的地步。
英文摘要
In an era where software development is deeply tied with Large Language Models, does Model Driven Engineering (MDE) still make sense? This raises the question of the extent to which MDE can be successfully combined with an LLM approach to address the downsides of each approach separately. In this paper, we try to answer that question by investigating a central Research Question: Can Agents powered by Large Language Models improve code generated from UML diagrams and text specifications by Model Driven Engineering (MDE) tools? To investigate this problem, we developed a novel LLM-powered Multi-Agentic approach called ARTHUR (Architecture Refactoring Through Hybrid UML Reasoning), a framework that aims at combining the reliability of MDE and the ease of use of LLMs. ARTHUR is designed to refactor legacy Java code produced by traditional, rule-based, MDE code generators from UML diagrams. To asses the answer to our question, we have refactored the legacy code of several projects from our custom dataset crafted for this purpose. \name{} which made it possible to add support for modern frameworks like Spring Boot, while ensuring compliance with Model-Based Testing techniques to verify that the code still corresponds to the initial model's specifications. We then measured the results obtained in terms of time and cost, passing test rate, and \texttt{compile@k}, \texttt{pass@k} and \texttt{pass$^k$} metrics. We also observed the effect of generating code directly from the conceptual model without the refactoring. Our preliminary test results show that MDE could not be more far from retirement, after all.