arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

EXplore Your Graphs ENgine (EXYGEN):实现知识图谱规模化理解

Enabling Knowledge Graph Understanding at Scale with the EXplore Your Graphs ENgine (EXYGEN)

Harshdeep Singh, Yurui Zhu, Giovanni Colavizza, Matteo Romanello

arXiv 2609.11569首次发表:更新:

发表机构

Odoma Ltd.(Odoma公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出EXYGEN框架,利用RAG和谓词覆盖感知采样实现知识图谱规模化对话式访问,无需微调即可在SciQA上达到0.419精确匹配,并将元数据生成速度提升80倍以上。

AI 中文摘要

我们提出了EXYGEN(EXplore Your Graphs ENgine),一个用于知识图谱(KG)理解的框架,能够实现对知识图谱的规模化对话式访问。我们依次解决两个问题。首先,在仅给定自动推导的结构化元数据和小规模图样本,而非针对特定任务进行微调的情况下,大型语言模型(LLM)执行文本到SPARQL生成的效率如何?我们将VoID描述和ShEx模式集成到检索增强生成(RAG)流水线中,并在SciQA基准上对知识图谱派生的上下文进行消融实验。我们最佳的配置——结合ShEx模式、检索到的三元组和示例问题-查询对——在执行结果上达到了0.419的精确匹配,且无需任何LLM微调。我们进一步发现,诸如F1之类的词汇指标难以预测查询的正确性,并且一旦提供足够的上下文,较大的通用LLM可以胜过较小的代码专用模型。其次,我们探讨如何从非常大的知识图谱中生成该方法所依赖的结构化元数据,因为在这些图谱上,知识图谱元数据的生成在计算上变得难以处理。我们引入了一种谓词覆盖感知的并行图采样策略,该策略在保持结构多样性的同时保持计算可行性。在OpenCitations Meta和GESIS上,它以最小的三元组损失保持了高谓词覆盖率,并将运行时间减少了80倍以上;在ORKG上,采样不仅更快,而且是获取完整元数据的唯一可行途径。总之,这些结果表明,结构化的模式上下文和轻量级提示可以大幅减少对微调的依赖,以实现可扩展的知识图谱对话式访问,尽管要弥合与完全微调方法之间的剩余差距,可能仍需减少对精心策划的问题-查询示例的依赖——无论是通过合成生成还是执行反馈驱动的方法——并在单个基准之外验证这些发现。

英文摘要

We present EXYGEN (EXplore Your Graphs ENgine), a framework for knowledge graph (KG) understanding that enables conversational access to KGs at scale. We address two questions in sequence. First, how effectively can LLMs perform text-to-SPARQL generation given only automatically derived structured metadata and small graph samples, rather than task-specific fine-tuning? We integrate VoID descriptions and ShEx schemas into a retrieval-augmented generation (RAG) pipeline and ablate KG-derived context on the SciQA benchmark. Our best configuration -- combining ShEx schemas, retrieved triples, and example question-query pairs -- reaches an exact match of 0.419 on execution results without any LLM fine-tuning. We further find that lexical metrics such as F1 poorly predict query correctness, and that larger general-purpose LLMs can outperform smaller code-specialized ones once given sufficient context. Second, we ask how to generate the structured metadata that this method relies on from very large KGs, where KG metadata generation becomes computationally intractable. We introduce a predicate-coverage-aware parallel graph sampling strategy that preserves structural diversity while remaining computationally tractable. On OpenCitations Meta and GESIS, it retains high predicate coverage with minimal triple loss and reduces runtime by over 80x; on ORKG, sampling is not just faster but the only tractable path to obtain complete metadata. Together, these results show that structured schema context and lightweight prompting can substantially reduce reliance on fine-tuning for scalable conversational access to KGs, though closing the remaining gap to fully fine-tuned approaches will likely require reducing dependence on curated question-query exemplars -- whether through synthetic generation or an execution-feedback-driven approach -- and validating these findings beyond a single benchmark.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑