arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.03729cs.CLcs.AIcs.DB

GPTKB 2.0:从大语言模型直接构建消歧知识库

Constructing Disambiguated Knowledge Bases from Large Language Models at Scale

Yujia Hu, Tuan-Phong Nguyen, Simon Razniewski

首次发表
浏览论文内容

中文总结 AI 辅助

针对从大语言模型构建知识库存在的重复与混淆问题,提出GPTKB 2.0方法,实现百万级消歧实体与三元组的知识库构建,为LLM原生知识库研究提供新方案。

中文摘要 AI 辅助

自动知识库构建(AKBC)是核心自然语言处理任务,近期研究提出从大语言模型(LLM)直接生成知识库,将模型本身作为知识源。但LLM原生不具备实体表示能力,会导致重复条目和实体混淆。我们提出GPTKB 2.0,一种从LLM直接构建消歧知识库的方法,该方法整合了实体、关系和类的实时消歧,且经过精心设计以兼顾可扩展性和消歧准确性。我们分析了核心设计决策,明确了准确性、规模与成本间的权衡关系。我们大规模部署GPTKB 2.0,得到了包含超100万个消歧实体和3840万个三元组的物化知识库,这是首个具备实体、关系和类显式内部规范化的百万级LLM原生知识库,与此前以维基媒体为中心的研究有显著区别。GPTKB 2.0可在指定网址获取。

英文摘要

Automated Knowledge Base Construction (AKBC) is a core NLP task, and recent work proposes generating knowledge bases directly from large language models (LLMs), treating the model itself as the knowledge source. However, LLMs natively possess no representation of entities, leading to duplicate entries as well as conflations. We propose GPTKB 2.0, a methodology for constructing disambiguated KBs directly from LLMs. GPTKB 2.0 incorporates on-the-fly disambiguation of entities, relations and classes, and is meticulously designed to satisfy both scalability and disambiguation accuracy. We analyze the central design decisions and characterize the trade-offs between accuracy, scale, and cost. We execute GPTKB 2.0 at scale, obtaining a materialized KB containing over 1M disambiguated entities and 38.4M triples. This represents the first million-scale LLM-native KB with explicit internal canonicalization of entities, relations, and classes, a significant departure from prior Wikimedia-centric works. GPTKB 2.0 is available at https://gptkb.org/.

补充信息

↑