从平面文化遗产记录自动构建FAIR数字对象知识图谱
Automated Construction of FAIR Digital Object Knowledge Graphs from Flat Cultural Heritage Records
- Fraunhofer Institute for Applied Information Technology FIT(弗劳恩霍夫应用信息技术研究所FIT)
- University Hospital of Cologne(科隆大学医院)
- University of Cologne(科隆大学)
- Foundation for Research and Technology(希腊研究与技术基金会)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
该研究提出了基于大语言模型的管道,将平面欧洲数据模型记录转换为符合FDO规范的知识图谱,实现了高比例元数据链接与跨语言实体合并,提升了文化遗产数据的可机器操作性。
AI中文摘要:
FAIR数字对象(FDO)框架要求元数据属性值尽可能以持久标识符(PID)的形式表达,以生成完全可被机器操作的图谱,其中每个引用都可解析。欧洲数据模型(Europeana Data Model)的设计早于FDO规范,它将大多数元数据值存储为纯文本,这足以满足人类浏览需求,但无法为自动化智能体提供跨记录或跨集合的可追踪信息。本文提出了一条管道,用于将平面欧洲数据模型记录转换为符合FDO规范、以CIDOC-CRM构建的知识图谱。根据FDO规范,我们将每个文化遗产实体建模为具有自身PID、类型、概要和元数据层的离散FDO。核心技术挑战在于自动区分FDO规定中必须成为PID引用(可解析实体)的值和可保留为文字(如注释、测量值、日期等终端叶子)的值。我们使用一个大语言模型(LLM)解决该问题,该模型对每个元数据值进行分类,将其路由至受控词汇表(Getty AAT、Wikidata、VIAF、PeriodO),并将其链接至共享实体FDO。我们使用来自5个欧洲数据模型提供者的637条考古记录进行评估,用LLM处理每条记录。该管道链接了86%的元数据槽,解析了欧洲数据模型未已丰富的58.5%的值;还合并了字节级匹配无法区分的跨语言表面形式,人工审查显示33次此类合并中有17次正确。图谱连通性并非其与字符串匹配的区别,FDO图谱的独特之处在于每个节点都有类型且可解析。
英文摘要:
The FAIR Digital Object (FDO) framework mandates that metadata attribute values be expressed as persistent identifiers (PIDs) wherever possible, to produce a fully machine-actionable graph in which every reference is resolvable. The Europeana Data Model was designed long before the FDO specification, and it stores most metadata values as plain text. This serves human browsing well enough, but gives an automated agent nothing to follow across records or collections. We present a pipeline that transforms flat Europeana records into an FDO-compliant knowledge graph structured with CIDOC-CRM. Following the FDO specification, we model every heritage entity as a discrete FDO with its own PID, type, profile, and metadata layer. The core technical challenge is automating the FDO-prescribed distinction between values that must become PID references (resolvable entities) and those that may remain literals (terminal leaves such as notes, measurements, and dates). We address this with a large language model that classifies each metadata value, routes it to a controlled vocabulary (Getty AAT, Wikidata, VIAF, PeriodO), and links it to a shared entity FDO. We evaluate using 637 archaeological records from five Europeana providers, processing each with the LLM. The pipeline links 86% of metadata slots, resolving 58.5% of values Europeana had not already enriched. It also merges cross-lingual surface forms that byte-identical matching keeps apart, where 17 of 33 such merges are correct on manual review. Graph connectivity does not separate this from string matching; what distinguishes the FDO graph is that every node is typed and resolvable.