发表机构
Max Planck Institute for Biological Cybernetics(马克斯·普朗克生物信息学研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对特定领域档案元数据查询难题,提出NLKGQ系统,借助OWL本体和LLMs,能零样本生成准确结构化查询。经神经影像档案元数据实验,最佳配置准确率达100%,明确影响准确率关键因素,还对比SPARQL与SQL,凸显OWL优势。
AI 中文摘要
研究人员需要回答关于特定领域档案内容的临时问题,但往往缺乏对元数据编写结构化查询的专业知识。研究表明,当在精心设计的网络本体语言(OWL)本体中捕获领域词汇和语义时,大语言模型(LLMs)可以在零样本情况下生成准确的结构化查询,无需微调、检索增强或多智能体编排。本文提出了自然语言知识图谱查询(NLKGQ)系统,这是一个框架和开发过程,可实现对此类档案中自然语言元数据的访问。该框架包括一个网络界面,帮助研究人员提出自然语言问题,由领域无关的工具通过大语言模型将其转换为SPARQL并针对知识图谱执行。开发过程首先在正式的OWL本体中捕获领域词汇和语义。特定领域代码然后从档案源中提取元数据并将其导入由本体定义的知识图谱。两者都设计为可跨领域重用。研究人员在来自大规模神经影像研究档案的元数据上演示了该系统,评估了多个大语言模型和本体表示。最佳配置在与领域专家共同开发的能力和回归问题集上实现了100%的准确率。对八种本体表示的消融研究表明,可读的实体名称和语义注释是影响准确率的主要因素,比模型选择或提示工程更重要。研究人员还将SPARQL与自动生成的SQL数据库作为查询后端进行了比较,表明OWL的结构特征在大语言模型驱动的查询生成方面比SQL DDL具有显著优势。本文的演示领域还需要在适度的机构硬件上使用本地大语言模型来解决人类受试者数据的隐私问题。
英文摘要
Researchers need to answer ad-hoc questions about the contents of domain-specific archives but often lack the expertise to write structured queries on the metadata. We show that when domain vocabulary and semantics are captured in a well-designed Web Ontology Language (OWL) ontology, Large Language Models (LLMs) can generate accurate structured queries zero-shot, without task-specific fine-tuning, retrieval augmentation, or multi-agent orchestration. We present the Natural Language Knowledge Graph Query (NLKGQ) system, a framework and development process that enables natural language access to metadata in such archives. The framework includes a web interface that helps researchers pose natural language questions, which a domain-agnostic harness translates to SPARQL via an LLM and executes against a knowledge graph. The development process begins with capturing domain vocabulary and semantics in a formal OWL ontology. Domain-specific code then extracts metadata from archive sources and imports it into a knowledge graph defined by the ontology. Both are designed for reuse across domains. We demonstrate the system on metadata derived from a large-scale neuroimaging research archive, evaluating multiple LLMs and ontology representations. The best configurations achieve 100% accuracy on a 21-question competency and regression test set developed with domain experts. An ablation study across eight ontology representations reveals that readable entity names and semantic annotations are the dominant factors in accuracy, more significant than model choice or prompt engineering. We also compare SPARQL to an auto-generated SQL database as query backends, showing that OWL's structural features provide a substantial advantage over SQL DDL for LLM-driven query generation. Our demonstration domain requires local LLMs on modest institutional hardware to address privacy concerns for human subject data.