发表机构
University of Southern California(南加州大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对大型语言模型生成连贯多图SysML模型能力不足的问题,提出SEMAADB数据集与基准,含3000个上下文和15000个图,评估显示语义修复与跨图一致性仍是挑战。
AI 中文摘要
系统工程师使用多种图来描述系统的结构和行为。工程师们共同创建这些图,以确保它们使用相同的元素并彼此保持一致。大型语言模型可以将图生成为文本或代码,这使得自动创建系统图成为可能。然而,它们生成连贯图集的能力尚未得到充分理解,现有的数据集和基准并未直接在大规模上衡量这一能力。我们引入了SEMAADB(系统建模工程助手AI数据集与基准),一个包含3,000个工程上下文和15,000个图的数据集。每个上下文包含五个相互关联的SysML视图:需求图、块定义图、活动图、状态机图和序列图。这里,视图是呈现系统某一方面的图。我们检查了图集的一致性和有效渲染。其中100个上下文的集合也经过了人工验证,并构成了基准测试集。我们在两个任务上评估了三个语言模型。在图修复中,最强的模型修复了64.3%的语义错误。在跨图更新中,当模型在相关图之间应用一个变更时,最佳传播F1分数为80.7%。结果表明,语法修复几乎已解决,但语义修复和图间一致性对模型而言仍是具有挑战性的任务。因此,SEMAADB既提供了大型图资源,也提供了一套用于衡量连贯多图SysML生成的基准。
英文摘要
Systems engineers use several diagrams to describe the structure and behavior of systems. Engineers create these diagrams together to make sure that they use the same elements and remain consistent with one another. Large language models can generate diagrams as text or code, which makes it possible to create system diagrams automatically. However, their ability to generate coherent sets of diagrams is not well understood, and existing datasets and benchmarks do not directly measure this ability at scale. We introduce SEMAADB (Systems Engineering Modeling Assistant with AI Dataset and Benchmark), a dataset of 3,000 engineering contexts and 15,000 diagrams. Each context contains five connected SysML views: Requirement, Block Definition, Activity, State Machine, and Sequence. Here, a view is a diagram that presents one aspect of a system. We checked the diagram sets for consistency and valid rendering. A set of 100 contexts is also human-verified and forms the benchmark test set. We evaluate three language models on two tasks. In diagram repair, the strongest model repairs 64.3% of semantic errors . In cross-diagram update the best propagation F1 is 80.7% when a model applies one change across related diagrams. The results show that syntax repair is nearly solved, but semantic repair and consistency across diagrams are still challenging tasks for models. SEMAADB therefore provides both a large diagram resource and a set of benchmarks for measuring coherent multi-diagram SysML generation.