arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.13257cs.SE

大语言模型能否学习并应用多层次建模语义?首次实证研究

Can LLMs Learn and Apply Multi-Level Modelling Semantics? A First Empirical Study

Yuhong Fu, Weixing Zhang, Bowen Jiang, Haowei Cheng, Karamjti Kaur, Markus Stumptner

首次发表
浏览论文内容

中文总结 AI 辅助

研究大语言模型能否学习应用多层次建模语义,通过让三种商业大语言模型在六种提示策略下为挑战生成模型并与参考对比,发现句法正确性可及,语义正确性部分实现,明确能力边界,为工业5.0建模工作流设计提供参考。

中文摘要 AI 辅助

工业5.0强调以人为本的工业系统设计,对建模工具提出了更高要求。多层次建模(MLM)能直接表示三个或更多抽象层次,但语义约束更复杂,模型正确性依赖于此。大语言模型(LLMs)在模型驱动工程中得到越来越多研究,但目前证据完全基于两级建模任务,其能否推广到语义不同的MLM尚待测试。本文首次对此问题进行实证研究。使用三种商业大语言模型在六种提示策略下为MULTI仓库挑战生成多层次模型,与手动验证的参考模型对比。结果表明句法正确性可及,但语义正确性仅部分实现,提示策略在精度和完整性间权衡,自检主要起规则检查作用。克劳德表现最平衡。这些结果明确了当前大语言模型在MLM方面的能力边界,为工业5.0以人为本、人工智能辅助建模工作流程的设计提供了参考。

英文摘要

Industry 5.0 emphasises human-centric industrial system design, placing additional demands on modelling tools. Multi-level modelling (MLM) can directly represent three or more abstraction levels, but this comes at the cost of more complex semantic constraints that model correctness depends on. Large Language Models (LLMs) have been increasingly studied in model-driven engineering, but this evidence rests entirely on two-level modelling tasks, and whether it generalises to MLM, whose semantics differ in kind, remains untested. This paper presents the first empirical study of this question. We have three commercial LLMs (GPT-5.4, Claude Opus 4.6, and Gemini 3.1 Pro) generate multi-level models for the MULTI Warehouse Challenge in the SLICER language under six prompting strategies, yielding 90 generated models compared against a manually validated reference using fourteen metrics. Syntactic correctness is within reach, but semantic correctness is only partially achieved, with Instantiation/Specialisation Correctness ranging from 52% to 79%. Models reproduce content stated explicitly in the task text, but rarely complete structure and constraints the text implies without stating. Prompting strategies trade off precision against completeness, and self-checking functions mainly as a rule checker rather than reliably improving alignment with the reference design. Among the three LLMs, Claude shows the most balanced profile. These results clarify the boundaries of current LLM capability for MLM and inform the design of human-centred, AI-assisted modelling workflows for Industry 5.0.

↑