arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AAS-RAIL:通过检索增强的上下文学习改进资产管理壳的信息提取

AAS-RAIL: Improving Information Extraction for Asset Administration Shells through Retrieval-Augmented In-Context Learning

Janek Groß, Jens Heidrich

arXiv 2609.07334首次发表:更新:

发表机构

University of Applied Sciences Mainz(美因茨应用科学大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对从PDF数据表生成资产管理壳的困难,提出检索增强上下文学习(RAIL)方法,动态选择公司特定示例,无需微调,在多种LLM上相对提升提取质量30.4%-52.4%。

AI 中文摘要

资产管理壳(AAS)是工业4.0和数字产品护照的基石,提供工业资产的标准化数字表示。尽管制造商已经维护了广泛的技术产品文档,但从现有产品数据表中生成AAS实例仍然是一项劳动密集型任务,因为技术信息是从异构文档结构中提取的,并且通常涉及公司特定的术语和惯例。在这项工作中,我们提出了AAS-RAIL,一种检索增强的信息提取(IE)方法,使用大型语言模型(LLMs)从PDF产品数据表自动生成资产管理壳。所提出的检索增强上下文学习(RAIL)方法不依赖于固定的少样本示例集,而是从相似的资产管理壳中检索LLM生成的提取辅助工具,以提供实例特定的上下文学习(ICL)。这使得模型能够适应公司特定的命名约定和格式风格,而无需微调。我们的核心贡献是为每个数据表动态选择公司特定的AAS示例,用适应实例并组合语义检索和结构化信息提取的提取管道取代静态提示。所提出的方法在一组工业产品数据表上使用一系列开放和封闭权重的LLM进行了评估。实验结果表明,RAIL始终优于传统的少样本提示,提取质量相对提高了30.4%至52.4%。这些结果证明,我们的方法为公司特定的AAS生成提供了有效的改进。

英文摘要

The Asset Administration Shell (AAS) is a cornerstone of Industry 4.0 and the Digital Product Passport, providing standardized digital representations of industrial assets. While manufacturers already maintain extensive technical product documentation, generating AAS instances from existing product datasheets remains a labor-intensive task because technical information is extracted from heterogeneous document structures and often involves company-specific terminology and conventions. In this work, we present AAS-RAIL, a retrieval-augmented information extraction (IE) approach that automatically generates Asset Administration Shells from PDF product datasheets using large language models (LLMs). Instead of relying on a fixed set of few-shot examples, the proposed retrieval-augmented in-context learning (RAIL) approach retrieves LLM-generated extraction helpers from similar Asset Administration Shells to provide instance-specific in-context learning (ICL). This enables the model to adapt its extraction behavior to company-specific naming conventions and formatting styles without fine-tuning. Our core contribution is the dynamic selection of company-specific AAS examples for each datasheet, replacing static prompting with an extraction pipeline that adapts to instances and combines semantic retrieval and structured information extraction. The proposed approach is evaluated on a collection of industrial product datasheets using a selection of open- and closed-weight LLMs. Experimental results show that RAIL consistently improves extraction quality over conventional few-shot prompting, yielding relative improvements of 30.4-52.4%. These results demonstrate that our approach provides an effective improvement for company-specific AAS generation.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑