arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

连接前的整理:生产型知识图谱中的身份与本体标注

Curate Before You Connect: Identity and Ontology Tagging in a Production Knowledge Graph

Vaibhav Dangaich, Kevin Lewis, Kundeshwar Pundalik

arXiv 2608.10644首次发表:更新:

发表机构

Konectu(科内克图)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文针对生产型知识图谱的身份与本体标注问题,提出记录身份阶梯策略及锚定证据的本体标注方法,构建了含53万余实体的政府文档知识图谱,解决了实体合并的不可逆转错误等问题。

AI 中文摘要

信息抽取会生成候选实体与关系,将它们写入图谱的环节决定了实体的身份,而身份决策的破坏性是信息抽取错误所不具备的——错误的类型后续可以修正,但两个记录在同一身份下合并后,一旦属性被合并就无法拆分,且合并后不会留下错误痕迹。本文描述了摄入与本体标注层,该层将经过验证的抽取流转化为包含537157个实体、2198567个关系的知识图谱,这些数据来自98795份政府文档。我们提出了记录身份阶梯,该阶梯通过标识符列、名称列、显示名称和类型范围内的位置而非名称相似度来判定同一性;该阶梯管控解析后表格内的去重,而图谱写入环节采用更粗略的规范名称键,因此共享同一规范名称的记录会在完全相等时自动合并。我们认为自动化的边界应设在此处(而非通过演示证明):未报告任何身份基准,且该键允许的过度合并在构建层面无法被检测。此策略下实体解析仅标记候选,后续发生过一起事件:同一名称的两种表面形式被合并,损坏了一条正确记录并删除了来自无关文档的8个实体,该策略正是由此事件确立。随后我们描述了多类本体标注及一个未预料到的证据不对称性:实体名称是实例标签而非类型断言,因此将名称片段与类索引匹配会生成分类;要求锚定证据后,从36个减少到4个角色分配,且全部经确认正确。我们量化了图谱的一致性债务,展示了次级分类如何补偿父类归属错误的主类,并描述了整理队列已增长到48403条待处理提案,对应775个人类决策。

英文摘要

Extraction produces candidate entities and relationships; writing them into a graph is where identity is decided, and identity decisions are destructive in a way extraction errors are not. A wrong type can be corrected later, but two records merged under one identity cannot be separated once their properties have been combined, and the merge leaves no error behind. This paper describes the ingestion and ontology-tagging layer that turns a validated extraction stream into a knowledge graph of 537,157 entities and 2,198,567 relationships drawn from 98,795 government documents. We describe a record-identity ladder that decides sameness from identifier columns, name columns, display names and type-scoped position rather than from name similarity. The ladder governs de-duplication within parsed tables, while the graph write applies a coarser canonical-name key, so records sharing a canonical name merge automatically on exact equality. We argue rather than demonstrate that this is where the automation line belongs: no identity benchmark is reported, and the over-merges the key permits are undetectable by construction. That policy, under which entity resolution only ever flags candidates, followed an incident in which two surface forms of one name were merged, corrupting a correct record and deleting eight entities from an unrelated document. We then describe multi-class ontology tagging and an evidence asymmetry we did not anticipate: an entity name is an instance label rather than a type assertion, so matching name fragments against a class index invents classifications. Requiring anchored evidence cut role assignments on an enriched sample from 36 to 4, all confirmed correct. We quantify the graph's conformance debt, show secondary classifications compensating for a mis-parented primary class, and describe a curation queue grown to 48,403 pending proposals against 775 human decisions.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑