arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

半结构化数据自动化知识图谱构建的基准评测

Benchmarking Automated Knowledge Graph Construction from Semi-Structured Data

Tarek Al Mustafa, Birgitta König-Ries

arXiv 2609.26985首次发表:更新:

发表机构

German Centre for Integrative Biodiversity Research Halle-Jena-Leipzig – iDiv; Friedrich Schiller University Jena(德国综合生物多样性研究中心哈雷-耶拿-莱比锡; 耶拿弗里德里希·席勒大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对半结构化数据构建知识图谱缺乏全面评估的问题,提出包含六个质量维度的基准与评估流水线,提供十个数据集和指标套件,并用两个参考系统验证。该工作扩展了现有评估框架,兼顾映射预测与RDF数据生成及下游使用评估。

AI 中文摘要

知识图谱(KGs)在众多应用中发挥着越来越重要的作用,其应用范围从传统的知识表示扩展到作为大语言模型(LLMs)的记忆以支持下游任务。然而,知识图谱的构建是劳动密集型的;因此,近年来,人们提出了许多自动化该过程的方法。与专注于文本输入数据的构建方法的普及程度相比,针对半结构化输入的方法仍然代表性不足,因此,目前尚无全面的基准和评估套件来判断映射预测和生成的知识图谱的质量。这是有问题的,因为知识图谱的质量直接影响其支持的下游应用,因此,迫切需要为其构建建立强大的评估机制。在这项工作中,我们因此聚焦于从半结构化数据构建知识图谱的评估,并提出了一个用于知识图谱构建的基准和评估流水线,该流水线结合了以下质量维度:(1)语法有效性,(2)语义准确性,(3)一致性,(4)简洁性,(5)完整性,以及(6)基于知识图谱回答能力问题能力的实用性质量。这项工作贡献了一个现实的任务定义,扩展了当前最先进的评估框架,允许评估预测映射和RDF数据的系统,并同时评估知识图谱构建和下游使用。我们提供了一个全面的指标套件,提供了来自七个领域的十个专家策划的数据集,并使用两个参考系统展示了评估。

英文摘要

Knowledge Graphs (KGs) play an increasingly important role in numerous applications ranging from traditional knowledge representation to serving as memory for LLMs to support downstream tasks. However, their construction is labor-intensive; thus, in recent years, numerous approaches for automizing this process have been proposed. Compared to the popularity of construction approaches that focus on textual input data, methods for semi-structured inputs remain underrepresented and as a result, no comprehensive benchmark and evaluation suite exists to judge the quality of mapping predictions and generated KGs. This is problematic, as a KG's quality has direct influence on the downstream applications it supports and thus, strong evaluation mechanisms for their construction are urgently needed. In this work, we thus focus on the evaluation of KG construction from semi-structured data and present a benchmark and evaluation pipeline for KG construction that combines the quality dimensions (1) syntactic validity, (2) semantic accuracy, (3) consistency, (4) conciseness, (5) completeness, and (6) pragmatic quality measured on a KG's ability to provide answers to competency questions. This work contributes a realistic task definition, extends current state of the art evaluation frameworks, allows evaluation of systems that predict mappings and RDF data alike, and combines evaluation of both KG construction and downstream usage. We provide a comprehensive metrics suite, provide ten expert-curated datasets from seven domains, and showcase evaluation using two reference systems.

Commentspreprint - work in process document

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑