从查询到叙事:面向知识图谱探索与质量评估的文化遗产数据故事
From Queries to Narratives: Cultural Heritage Data Stories for Knowledge Graph Exploration and Quality Assessment
浏览论文内容
中文总结 AI 辅助
本文提出数据故事方法,通过将SPARQL查询与叙事文档结合,降低文化遗产知识图谱的探索门槛,同时实现数据质量评估,并以LODEON平台作为概念验证。
中文摘要 AI 辅助
文化遗产知识图谱(如NFDI4Culture-KG)包含数百万条关于艺术品、音乐、铭文、历史事件及其相关人物和地点的三元组。然而,对许多用户而言,发现这些知识可能很困难。虽然SPARQL可以学习,但编写有意义的查询首先需要深入理解图的数据模型,这是许多领域研究人员和实践者不愿投入的。即使有现有的用户界面,通常也需要一个起点和一些指导,因为图中包含的数据高度专业化、异构且不断增长,使得了解其内容或它能回答哪些问题变得具有挑战性。在本文中,我们提出数据故事作为一种方式,不仅降低这一障碍,还将探索转变为数据质量评估,从而将可访问的查询与发现隐藏在聚合统计中的问题结合起来。在此贡献中,数据故事被理解为一种叙事文档,将解释性文本和图像与可执行的SPARQL查询及其可视化结果相结合。我们描述了如何针对图编写这些故事,以及它们如何服务于多个目的:引导用户浏览不熟悉的图、创建可复现的叙事、以及揭示先前隐藏在聚合统计中的数据质量问题。作者平台LODEON(包括其Sparnatural和AI支持的作者助手)作为概念验证被引入。在作者环境中,关于数据的每个声明都可以由显式查询支持,使这些叙事透明且可复现。本文还反思了从实践研讨会和工作坊中获得的经验教训。早期经验表明,此类数据故事使文化遗产知识图谱在探索和质量评估方面都更易于访问。
英文摘要
Cultural-heritage KGs such as the NFDI4Culture-KG contain millions of triples about artworks, music, inscriptions, historical events, and the people and places connected to them. For many users, however, discovering this knowledge can be difficult. While SPARQL can be learned, writing meaningful queries first requires an in-depth understanding of the graph's data model, an investment many domain researchers and practitioners are unwilling to make. Even with existing user interfaces, a starting point and some guidance are usually needed, because the data contained in the graph is highly specialized, heterogeneous, and constantly growing, making it challenging to know what it contains or which questions it can answer. In this paper, we present data stories as a way not only to lower this barrier, but also to turn exploration into data-quality assessment, and thus combine accessible querying with the discovery of issues that remain hidden in aggregate statistics. In this contribution, a data story is understood as a narrative document that integrates explanatory text and images with executable SPARQL queries and their visualized results. It is described how they are authored against the graph and how they serve several purposes: guiding users through an unfamiliar graph, creating reproducible narratives, and surfacing data-quality issues previously hidden in aggregate statistics. The authoring platform LODEON including its Sparnatural and AI-supported authoring assistants is introduced as a proof-of-concept. Within the authoring environment, every claim made about the data can be backed by an explicit query, making these narratives transparent and reproducible. This paper also reflects on lessons learned from hands-on seminars and workshops. Early experience suggests that such data stories make cultural-heritage knowledge graphs more accessible for both exploration and quality assessment.
发表机构
- FIZ Karlsruhe – Leibniz Institute for Information Infrastructure(FIZ卡尔斯鲁厄–莱布尼茨信息基础设施研究所)
- Institute of Applied Informatics and Formal Description Methods (AIFB) of KIT(卡尔斯鲁厄理工学院应用信息学与形式描述方法研究所)
- Academy of Sciences and Literature Mainz(美因茨科学与文学院)
机构由 AI 辅助整理,请以论文原文为准。