发表机构
Télécom Paris; Institut Polytechnique de Paris(巴黎高等电信学院; 巴黎理工学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对Wikidata的类型不一致问题,提出WiCleanData,通过语言模型辅助分类清理、层次聚合简化约束及事实过滤,构建无类型违反的知识图谱并公开。
AI 中文摘要
由于其协作性质,Wikidata存在错误、不一致和过度复杂的问题,例如冗余类、实例与类之间的歧义、错误的分类路径以及类型约束违反。手动整理这些问题在大规模上不可行。为应对这些挑战,我们引入了WiCleanData,这是Wikidata的一个精炼版本,具有一致的分类法且无类型约束违反。具体而言,我们设计了一个自动化流水线,首先借助语言模型辅助清理分类法,然后通过层次聚合简化类型约束,最后据此过滤事实。由此产生的知识图谱无任何类型违反,并通过Web界面公开提供,便于探索和下游应用。
英文摘要
Because of its collaborative nature, Wikidata suffers from errors, in- consistencies, and excessive complexity, such as redundant classes, ambiguity between instances and classes, wrong taxonomic paths, and type constraint violations. The manual curation of these issues is infeasible at scale. To address these challenges, we introduce WiCleanData, a refined version of Wikidata with a consistent tax- onomy and free from type constraint violations. Specifically, we have designed an automated pipeline that first cleans the taxonomy with language model assistance, then simplifies type constraints by hierarchical aggregation, and finally filters facts accordingly. The resulting knowledge graph, free from any type violation, is made publicly available via a Web interface, enabling easy exploration and downstream applications.
Journal refCIKM, 2026, Rome, Italy