arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

KG2Code:通过可执行代码连接知识图谱与大语言模型以进行问答

KG2Code: Bridging Knowledge Graphs and Large Language Models via Executable Code for Question Answering

Yike Wu, Nan Hu, Guilin Qi, Guohui Xiao, Chen Jiang, Xinchun Zou, Yuchen Lu, Songlin Zhai, Yongrui Chen, Yuyang Zhang, Xiaoguang Li, Lifeng Shang, Jiaoyan Chen, Jeff Z. Pan

arXiv 2607.22652首次发表:更新:

发表机构

Southeast University; Key Laboratory of New Generation Artificial Intelligence Technology and Its Interdisciplinary Applications (Southeast University), Ministry of Education; Huawei Technologies; University of Manchester; University of Edinburgh(东南大学; 教育部新一代人工智能技术及其交叉应用重点实验室(东南大学); 华为技术有限公司; 曼彻斯特大学; 爱丁堡大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对知识图谱问答中现有方法的局限,提出KG2Code方法,将知识图谱转换为代码表示,在此基础上构建KG2Code-QA框架,把KGQA作为代码生成任务,还构建代码语料库,训练后的模型在KGQA任务中表现优异且泛化性强。

AI 中文摘要

近期研究探索将知识图谱(KGs)与大语言模型(LLMs)整合,以提升其在下游知识密集型任务上的表现,尤其是知识图谱问答(KGQA)。现有方法主要通过基于检索增强生成(RAG)、基于智能体和基于SPARQL的方法将LLMs与KGs结合。虽取得一定成功,但仍有结构信息丢失、推理不准确及灵活性和泛化性有限等局限。本文提出KG2Code,将知识图谱转换为基于代码的表示,保留结构语义并与现代LLMs的代码感知预训练自然对齐。在此基础上,KG2Code-QA框架将KGQA表述为代码生成任务,可生成可验证推理轨迹和可执行代码,减轻幻觉影响。还开发自动化管道构建大规模高质量代码语料库,训练后的LLMs能在零样本场景下执行KGQA。实验表明该方法显著优于现有KG增强的LLM方法,且对未见KGs有强泛化性。代码和数据可在Github获取。

英文摘要

Recent research has explored the integration of knowledge graphs (KGs) with large language models (LLMs) to enhance their performance on downstream knowledge-intensive tasks, particularly knowledge graph question answering (KGQA). Existing approaches primarily combine LLMs with KGs through retrieval-augmented generation (RAG)-based, agent-based, and SPARQL-based methods. Although these methods have achieved notable success, they still suffer from several limitations, including structural information loss, unfaithful reasoning, and limited flexibility and generalization. To address these challenges, this paper proposes KG2Code, a novel approach that transforms knowledge graphs into a code-based representation, preserving structural semantics while naturally aligning with the code-aware pretraining of modern LLMs. Based on KG2Code, KG2Code-QA is further introduced as a KGQA framework that formulates KGQA as a code generation task. This formulation enables the generation of verifiable reasoning traces and executable code, thereby substantially mitigating the impact of hallucinations. In addition, an automated pipeline is developed to construct a large-scale, high-quality code corpus for effectively training open-source LLMs on KG2Code-QA. After training, LLMs are able to perform KGQA in zero-shot scenarios. Extensive experiments demonstrate that the proposed approach significantly outperforms existing KG-enhanced LLM methods for KGQA, while exhibiting strong generalization to unseen KGs. The code and data are available at Github.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑