发表机构
Max Planck Institute for Biological Cybernetics(马克斯·普朗克生物控制论研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
NLKGQ通过OWL本体在LLM上下文中传递受控语义,零样本生成SPARQL查询,并在多个知识图谱基准上显著提升查询匹配率。
AI 中文摘要
大型语言模型(LLM)应用通常通过提示文本、模式转储和示例,以非正式方式将领域概念传递到模型的上下文中。我们表明,对于数据库查询,数据模型概念通过其词汇项带有声明的、机器可读语义(受控语义)的表示形式,能更有效地传递给LLM。NLKGQ是一个可工作的系统和可复用框架,用于对以知识图谱建模的数据执行此操作。一个正式的OWL本体作为传递机制,将数据的含义集中到模型可直接使用的语义精确的标记中。在单次LLM调用中,NLKGQ将系统提示(指示SPARQL)、完整的领域OWL本体、领域特定的提示补充以及用户的自然语言查询放入上下文中。然后模型直接零样本生成SPARQL查询。当现有数据库或数据库联合的本机词汇不透明时,包装本体替换为清晰的术语,运行时重写器恢复本机形式。在DBLP-QuAD 2.0上的评估表明,其得分取决于图谱快照、使用的端点和机器生成问题的措辞,因此我们提出DBLP-QuAD 3.1,该基准在保持2.0意图的同时,使参考结果确定性化,在需要时修订参考SPARQL,并使用前沿模型重写自然语言问题,以清晰完整地陈述每个参考查询的意图。我们在DBLP-QuAD 2.0基准(确定性重新评分下匹配率为57.6%)、DBLP-QuAD 3.1(1000个问题上匹配率为89.9%)、SemOpenAlex(相同测试集上匹配率为98%,而已发布基线的匹配率为86%)以及神经影像元数据(100%)上进行了评估。
英文摘要
Large Language Model (LLM) applications often transfer domain concepts into the model's context informally, through prompt prose, schema dumps, and examples. We show that for database queries, data model concepts pass to LLMs more effectively through representations whose vocabulary terms carry declared, machine-readable semantics (controlled semantics). NLKGQ is a working system and reusable framework that does this for data modeled in a knowledge graph. A formal OWL ontology serves as the transfer mechanism, concentrating the meaning of the data into semantically precise tokens the model can use directly. In a single LLM call, NLKGQ places in the context a system prompt instructing on SPARQL, the complete domain OWL ontology, and a domain-specific prompt addition, together with the user's natural language query. The model then generates the SPARQL query directly, zero-shot. Where the native vocabulary of an existing database or federation of databases is opaque, a wrapper ontology substitutes clean terms and a runtime rewriter restores the native forms. Evaluating on DBLP-QuAD 2.0 showed that its scores depend on the graph snapshot, the endpoint used, and the wording of its machine-generated questions, so we propose DBLP-QuAD 3.1, which maintains the intent of 2.0 while making reference results deterministic, revising reference SPARQL where needed, and rewriting the natural language questions, with a frontier model, to state each reference query's intent clearly and completely. We evaluate on the DBLP-QuAD 2.0 benchmark (57.6% Match under deterministic re-scoring), DBLP-QuAD 3.1 (89.9% Match on 1,000 questions), SemOpenAlex (98% Match against a published baseline's 86% on the identical test set), and neuroimaging metadata (100%).
Comments9 pages, 4 tables