发表机构
ScaDS.AI Dresden/Leipzig; TU Dresden; Institute for AI, VNU University of Engineering and Technology(ScaDS.AI德累斯顿/莱比锡; 德累斯顿工业大学; 越南国家大学工程技术学院人工智能研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究推出了可浏览、查询、审计的GPTKB 2.0系统,其为经上下文引导消歧的LLM衍生知识库,规模达3840万条三元组,支持多种查询及实体链接功能,且提供网络演示与离线下载。
AI 中文摘要
我们推出一款用于探索从大语言模型(LLM)生成的大规模经消歧知识库(KB)的网络演示系统。GPTKB 2.0包含160万个规范实体上的3840万条三元组,以及20.76万条整合关系和6.6万个整合类别。与此前主要通过表面字符串识别实体的LLM衍生知识库不同,GPTKB 2.0在递归构建知识库过程中执行上下文引导的消歧,在提取事实时区分同形异义词并合并同义提及。该演示系统使这一过程可被检查:用户可浏览实体、跟踪知识库内的链接,以及审计单个事实的来源,包括表面形式、候选匹配、来源三元组和消歧决策。该界面还支持结构化SPARQL查询、转换为SPARQL的自然语言问题,以及将用户提供的文本中的实体链接到GPTKB 2.0的规范条目。GPTKB 2.0可通过此URL访问,完整知识库可下载用于离线使用。
英文摘要
We present a web demo for exploring a large-scale disambiguated knowledge base (KB) materialized from a large language model (LLM). GPTKB 2.0 contains 38.4M triples over 1.6M canonical entities, together with 207.6K consolidated relations and 66K consolidated classes. Unlike prior LLM-derived knowledge bases that largely identify entities by surface strings, GPTKB 2.0 performs context-guided disambiguation during recursive KB construction, separating homonyms and merging synonymous mentions as facts are elicited. The demo makes this process inspectable: users can browse entities, follow links across the KB, and audit the provenance of individual facts, including surface forms, candidate matches, source triples, and disambiguation decisions. The interface further supports structured SPARQL queries, natural-language questions translated to SPARQL, and entity linking from user-provided text to canonical GPTKB 2.0 entries. GPTKB 2.0 is available at https://gptkb.org/, with the full KB downloadable for offline use.
CommentsAccepted to EMNLP 2026 Demo Track