发表机构
University of Applied Sciences Mainz(美因茨应用科学大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出基于扰动的评估框架,利用GPT-4o-mini、Qwen3和DeepSeek-R1生成6400个AAS实例,发现基于值召回率和名称F1分数的指标最能可靠反映LLM生成资产管理壳的质量退化,为工业应用中的质量保证提供支持。
AI 中文摘要
制造业的快速数字化转型,通常被称为工业4.0,依赖于物理资产和软件资产之间的无缝互操作性。一个关键的推动因素是资产管理壳(AAS),它是此类资产的标准化数字表示。大型语言模型(LLM)的最新进展使得能够从非结构化来源(如产品数据表)生成AAS子模型,但这给质量保证带来了挑战。特别是,意外错误、缺乏真实参考以及缺乏标准化质量指标阻碍了可靠采用。在这项工作中,我们使用基于扰动的评估框架来评估AI生成的AAS的质量指标。通过沿多个维度系统地降低AAS生成质量,我们评估了不同指标反映质量变化的程度。基于来自多个制造商的200个产品的数据集,我们使用GPT-4o-mini、Qwen3和DeepSeek-R1生成了6,400个AAS实例。我们的结果表明,基于属性名称精确匹配和属性值基于相似性的软匹配的指标,特别是基于值的召回率和基于名称的F1分数,提供了最可靠的质量退化指标。此外,我们量化了不同扰动类型的影响,并分析了不同模型系列和产品细分之间的差异。这些发现支持选择适当的指标、调整基于LLM的管道以及将AI生成的AAS集成到工业应用中。
英文摘要
The rapid digital transformation of manufacturing, often referred to as Industry 4.0, relies on seamless interoperability between physical and software assets. A central enabler is the Asset Administration Shell (AAS), a standardized digital representation of such assets. Recent advances in large language models (LLMs) enable the generation of AAS submodels from unstructured sources such as product datasheets but raise challenges for quality assurance. In particular, unexpected errors, the lack of ground truth references, and the absence of standardized quality metrics hinder reliable adoption. In this work, we evaluate quality metrics for AI-generated AAS using a perturbation-based evaluation framework. By systematically degrading AAS generation along multiple dimensions, we assess how well different metrics reflect quality changes. Based on a dataset of 200 products from multiple manufacturers, we generate 6,400 AAS instances using GPT-4o-mini, Qwen3, and DeepSeek-R1. Our results show that metrics based on exact matching of property names and similarity-based soft matching of property values, in particular value-based recall and name-based F1 score, provide the most reliable indicators of quality degradation. Furthermore, we quantify the impact of different perturbation types and analyze differences across model families and product segments. These findings support the selection of suitable metrics, the tuning of LLM-based pipelines, and the integration of AI-generated AAS into industrial applications.