arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.31663cs.CL

LLM引导的基于本体的非结构化文本知识图谱构建

LLM-Guided Ontology-Driven Knowledge Graph Construction from Unstructured Text

  • Institute for Technological Research - IRT SystemX(技术研究院——IRT SystemX)

机构由 AI 辅助整理,请以论文原文为准。

Abdelhadi Belfadel, Maxence Gagnant, Joseph Kattan, Sana Tmar

AI总结:

本文提出一种结合开源大语言模型、可复用提示策略和开放知识库的本体学习流水线,用于从非结构化工业文本构建知识图谱,并在法语电网事故报告上验证了其有效性与泛化能力。

AI中文摘要:

从工业文本中进行基于本体的知识图谱构建仍然具有挑战性,原因在于文档的领域特异性、标注资源的稀缺性以及本体工程工作流的复杂性。本文提出并研究了一种本体学习流水线的适用性,该流水线结合了紧凑型开源大语言模型(LLMs)、可复用的提示策略和开放知识库,以支持从文本语料库中提取、结构化、丰富和评估知识。该方法在私有的法语电网事故报告语料库上进行了测试和评估,使用了从7B到32B参数不等的本地可部署开源LLMs。该方法从非结构化报告出发,提取实体和关系,生成RDF三元组,构建相关的OWL本体,利用外部知识源对其进行丰富,评估本体的质量,随后构建一个基于所得本体模式的知识图谱。对80份人工标注的私有报告进行的实验表明,模式引导的提示显著提高了提取质量,而量化模型在性能和计算成本之间提供了有效的权衡。这些结果证明了使用本地部署的开源LLMs将特定领域的工业文本转化为基于本体的知识图谱的可行性,同时通过可复用的提示策略支持了提取过程的泛化。

英文摘要:

Ontology-driven knowledge graph construction from industrial text remains challenging due to the domain specificity of documents, the scarcity of annotated resources, and the complexity of ontology engineering workflows. This paper presents and investigates the applicability of an ontology learning pipeline that combines compact open-source Large Language Models (LLMs), reusable prompting strategies, and open knowledge bases to support the extraction, structuring, enrichment, and evaluation of knowledge from textual corpora. The approach is tested and evaluated on a private French corpus of power-grid incident reports, using locally deployable open-source LLMs ranging from 7B to 32B parameters. Starting from unstructured reports, the approach extracts entities and relations, generates RDF triples, constructs related OWL ontology, enriches it using external knowledge sources, assesses the quality of the ontology, and subsequently constructs a populated knowledge graph grounded in the resulting ontology schema. Experiments on 80 manually annotated private reports show that schema-guided prompting significantly improves extraction quality, while quantized models provide an effective trade-off between performance and computational cost. These results demonstrate the feasibility of transforming domain-specific industrial text into ontology-based knowledge graphs using locally deployed open-source LLMs, while supporting the generalization of the extraction process through reusable prompting strategies.

补充信息

↑