AI 中文总结
该系统综述分析64项2023-2025年研究,指出LLM在UML/ER建模中存在领域失衡、GPT依赖、评估不规范等问题,提出需标准化基准等改进方向。
AI 中文摘要
大语言模型(LLMs)正越来越多地应用于基于图的软件和数据建模。在各类建模符号中,UML(统一建模语言)是软件建模应用最广泛的符号,实体-关系(ER)图则是数据建模应用最广泛的符号。近期文献已探究了LLMs在图建模中的各类应用,但其有效性与局限性尚未得到充分讨论。本系统文献综述分析了2023至2025年间发表的64项研究,考察了图覆盖范围、建模任务、技术方法、评估实践及局限性。研究结果显示存在显著的集中模式与缺口:基于UML的软件建模占绝对主导,其中类图受关注最多,而行为图与数据建模领域仍关注不足;从自然语言生成图是核心研究方向,关于转换、质量保证及一致性检查的研究有限;基于GPT的模型被大量使用,引发了可复现性及供应商依赖的担忧;评估实践呈现异质性,采用了多样的指标与自定义数据集,基准复用度有限,且鲁棒性与统计显著性的报告不一致。常见局限性包括语义不准确、图元素幻觉、对提示表述的敏感性及可复现性约束。本综述是首次对基于LLM的图建模研究进行系统综合,强调了对标准化基准、更严格的评估协议、更广泛的图覆盖,以及提升语义可靠性与多视图一致性的技术的需求。
英文摘要
Large language models (LLMs) are increasingly applied to diagram-based software and data modelling. Among various modelling notations, UML and entity-relationship (ER) diagrams are the most widely adopted for software modelling and data modelling, respectively. Recent literature has investigated various applications of LLMs in diagram modelling; however, their effectiveness and limitations have not been extensively discussed. This systematic literature review analyses 64 studies published between 2023 and 2025, examining diagram coverage, modelling tasks, technical approaches, evaluation practices, and limitations. Our findings reveal significant concentration patterns and gaps. UML-based software modelling strongly dominates, with class diagrams receiving the most attention whilst behavioural diagrams and data modelling remain underrepresented. Diagram construction from natural language is the primary focus, with limited work on transformation, quality assurance, and consistency checking. GPT-based models are heavily prevalent, raising concerns about reproducibility and vendor dependence. Evaluation practices are heterogeneous, employing diverse metrics and custom datasets with limited benchmark reuse and inconsistent reporting of robustness and statistical significance. Common limitations include semantic inaccuracies, hallucinated diagram elements, sensitivity to prompt formulation, and reproducibility constraints. This survey provides the first systematic synthesis of LLM-based diagram modelling research, highlighting needs for standardised benchmarks, stronger evaluation protocols, broader diagram coverage, and techniques for improving semantic reliability and multi-view consistency.