arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

面向大规模语义表格解释与数据质量评估的可解释性表头中心框架

An Explainable Header-Centric Framework for Large-Scale Semantic Table Interpretation and Data Quality Assessment

Marcelo Valentim Silva, Hannes Herrmann, Valerie Maxville

arXiv 2610.10541首次发表:更新:

发表机构

BCB(BCB机构)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出以表头为中心的可解释框架,用于仅元数据场景下的列类型标注与数据质量评估,在含约12万表头列的多基准上验证,还提供KG映射路径及基准分歧诊断方法。

AI 中文摘要

知识图谱(KG)的质量不仅取决于下游图验证,还取决于集成前使用的表格元数据的质量。在仅使用元数据的语义表格解释(STI)中,当单元格值不可用、存在噪声或不合适时,列表头就成为可追溯KG准备的关键语义证据来源。我们提出了一种可解释的、以表头为中心的框架,用于仅使用元数据的列类型标注(CTA)和数据质量评估(DQA)。该框架利用精心整理的词汇资源将表头映射到39种可解释的FinalFormat类型,并通过SourceKeywords保留词级可追溯性。每种分配的类型都会基于数据质量问题(DQIs)分类法激活验证规则,生成诸如缺失数据、重复项、域违反、错误数据类型和时间不匹配等检测结果。这些检测结果被汇总为HeadersIQ,这是一种轻量级、无权重的数据源级质量指标。该框架在异构基准上进行了评估,包括UCI、布拉格、Kaggle、VizNet/Sato、SOTAB、T2Dv2以及SemTab 2024元数据到KG赛道,涵盖约120000个表头列。结果显示其在有噪声的真实元数据中具有广泛的实际覆盖范围,同时并行的KG映射路径支持与DBpedia及该http URL的对齐。在SemTab 2024元数据到KG赛道中,官方GT-strict评估结果一般,但盲诊断审计表明,许多不匹配反映了基准粒度、别名和本体选择的影响,而非完全不可信的以表头为中心的预测。我们将此审计作为分歧模式的诊断证据报告,而非修订后的基准性能。总体而言,本文提出了一种可重用的工作流,用于元数据驱动的语义标注、数据源级质量监控和面向KG的基准诊断。

英文摘要

Knowledge Graph (KG) quality depends not only on downstream graph validation, but also on the quality of tabular metadata used before integration. In metadata-only Semantic Table Interpretation (STI), where cell values are unavailable, noisy, or unsuitable, column headers become a critical source of semantic evidence for traceable KG preparation. We present an explainable, header-centric framework for metadata-only Column Type Annotation (CTA) and Data Quality Assessment (DQA). The framework maps headers to 39 interpretable FinalFormat types using curated lexical resources and preserves token-level traceability through SourceKeywords. Each assigned type activates validation rules based on a taxonomy of Data Quality Issues (DQIs), producing detections such as missing data, duplicates, domain violations, wrong data type, and temporal mismatch. These detections are aggregated into HeadersIQ, a lightweight, unweighted data source-level quality metric. The framework was evaluated across heterogeneous benchmarks, including UCI, Prague, Kaggle, VizNet/Sato, SOTAB, T2Dv2, and the SemTab 2024 Metadata-to-KG track, comprising around 120,000 header columns. The results show broad practical coverage across noisy real-world metadata, while a parallel KG-mapping pathway supports alignment to DBpedia and Schema.org. On the SemTab 2024 Metadata-to-KG track, the official GT-strict evaluation was modest. However, a blinded diagnostic audit indicates that many mismatches reflect benchmark granularity, aliasing, and ontology-selection effects rather than wholly implausible header-centric predictions. We report this audit as diagnostic evidence on disagreement patterns, not as revised benchmark performance. Overall, the paper presents a reusable workflow for metadata-driven semantic annotation, data source-level quality monitoring, and KG-oriented benchmark diagnosis.

Comments18 pages, 4 figures, Workshop on Quality of Knowledge Graphs at ESWC 2026, May 11, 2026, Dubrovnik, Croatia

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑