arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

KGCQual:一个用于评估从文本构建知识图谱质量的可解释框架

KGCQual: An Interpretable Framework for Evaluating the Knowledge Graph Construction Quality from Text

Nipun Misra, Vikranth Udandarao, Aanchal Gupta, Yogender Kumar, Manuj Mukherjee, Raghava Mutharaju

arXiv 2607.10212首次发表:更新:

发表机构

VIT Vellore; IIIT-Delhi; Indian Institute of Technology Palakkad(维洛尔理工学院; 德里信息技术学院; 印度理工学院帕拉卡德分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对自动构建知识图谱质量评估问题,提出KGCQual框架,通过实体级和关系级评估,衡量提取图谱与理想图谱的接近度,能识别现有指标遗漏问题,为比较构建方法提供可扩展、可解释框架,还通过实验验证了指标有效性。

AI 中文摘要

知识图谱越来越多地通过自动提取管道构建,但此类系统常引入虚假或不完整的三元组,降低下游性能。现有评估方法依赖特定任务指标或小规模人工验证,对提取图谱的结构和语义保真度洞察有限。我们提出一种用于内在知识图谱质量评估的新颖、可解释指标,衡量自动提取图谱与捕捉源文本关键名词短语、谓词关系和基本语言现象(如否定)的“理想”图谱的接近程度。我们的框架整合两个互补组件:实体级评估和关系级评估。我们在多个数据集上评估该指标并验证,它能可靠识别现有指标忽略的问题,为比较自动知识图谱构建方法提供了可扩展、模型无关且可解释的框架。

英文摘要

Knowledge Graphs (KGs) are increasingly constructed through automated extraction pipelines; however, such systems often introduce spurious or incomplete triples, which degrade downstream performance. Existing evaluation practices rely heavily on task-specific metrics or small-scale manual verification, offering limited insight into the structural and semantic fidelity of extracted graphs. We propose a novel, interpretable metric for intrinsic KG quality assessment that measures how closely an automatically extracted graph approximates an "ideal" graph capturing the key noun phrases, predicate relations, and basic linguistic phenomena such as negation expressed in the source text. Our framework integrates two complementary components: (1) an entity-level assessment that evaluates completeness, resolution quality, and connectivity, and (2) a relation-level assessment that judges predicate preservation and multiplicity using lexical similarity, dependency-parse alignment, and light-weight negation handling to ensure semantic faithfulness. We evaluate our metric across multiple state-of-the-art triple extraction systems and datasets, including WebNLG, TinyButMighty, and BenchIE, demonstrating that it reliably identifies omissions, redundancy, and structural deviations that existing metrics overlook. Our work offers a scalable, model-agnostic, and interpretable framework for comparing automated KG construction methods and provides a foundation for standardised evaluation. We further validate the metric through an ablation study isolating noun and verb components, and a downstream evaluation showing that KGCQual scores correlate significantly with link prediction performance on the same extracted KGs. The code repository is available at https://github.com/kracr/kg-quality-metric.

CommentsThis paper is under review at a conference

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑