arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

KGCaRe:结合自动知识图谱构建与大语言模型上下文检索的可解释复杂条件问答方法

KGCaRe: Explainable Complex Conditional Question Answering using Automatic Knowledge Graph Construction and Context Retrieval with LLMs

Ghanshyam Verma, Simanta Sarkar, Devishree Pillai, Hotaka Shiokawa, Yourong Xu, Fiona Veazey, Peter Hubbert, Hui Su, Paul Buitelaar

arXiv 2608.09779首次发表:更新:

发表机构

University of Galway; Fidelity Investments(戈尔韦大学; 富达投资)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对LLM与RAG处理复杂条件问答时在特定领域表现不佳的问题,提出KGCaRe混合方法,结合KG构建、迭代图遍历与神经检索,在多类LLM及数据集上优于现有基线。

AI 中文摘要

使用大语言模型(LLM)与检索增强生成(RAG)回答复杂条件问题仍是一项挑战,尤其在特定领域语境下,通用LLM与RAG往往表现不佳。我们假设,用从文档和知识图谱(KG)中提取的非结构化与结构化知识增强RAG,可提升这类任务的推理能力与答案准确率。为验证这一点,我们提出KGCaRe,一种结合神经检索与对LLM生成的KG进行符号推理的混合方法。KGCaRe采用多提示提取策略从文档构建KG并将其存储在图数据库中,同时将文档嵌入向量存储以实现神经检索。KGCaRe执行由LLM引导的创新迭代图遍历,以提取相关三元组、修剪无关信息;若初始遍历未提供生成答案的满意上下文,还会使用额外线索实体再次遍历图。从KG中以路径形式提取的相关三元组,以及语义检索到的文本段落,随后被输入定制的KGCaRe提示,以生成带解释的复杂条件问题答案。我们在两个复杂条件QA数据集上评估KGCaRe,结果显示,在Mistral、Mixtral、GPT-3.5、GPT-4o等多个LLM上,KGCaRe在包括Vanilla LLM、Code Prompt、Text Prompt、Think-on-Graph、Vanilla RAG、HybridContextQA在内的现有基线中始终表现更优。我们公开发布了为实现所提KGCaRe方法而开发的软件流水线。

英文摘要

Answering complex conditional questions using Large Language Models (LLMs) and Retrieval-Augmented Generation (RAG) remains a challenge, particularly in domain-specific contexts where general-purpose LLMs and RAG tend to underperform. We hypothesize that augmenting RAG with unstructured and structured knowledge, extracted from both documents and knowledge graphs (KGs), can improve reasoning and answer accuracy for such tasks. To test this, we propose KGCaRe, a hybrid approach that combines neural retrieval with symbolic reasoning over LLM-generated KGs. KGCaRe constructs a KG from documents using a multi-prompt extraction strategy and stores it in a graph database. Simultaneously, the documents are embedded into a vector store to enable neural retrieval. KGCaRe performs innovative iterative graph traversal guided by the LLM to extract relevant triples, prune irrelevant information, and uses additional clue entities to traverse the graph again if the initial traversal does not provide satisfactory context to generate the answer. The relevant triples extracted from the KG in path form, along with semantically retrieved text passages, are then fed into custom KGCaRe prompts to generate answers to the complex conditional questions with explanations. We evaluate KGCaRe on two complex conditional QA datasets. Our results on these datasets show that KGCaRe consistently outperforms existing baselines, including Vanilla LLM, Code Prompt, Text Prompt, Think-on-Graph, Vanilla RAG, and HybridContextQA, across multiple LLMs such as Mistral, Mixtral, GPT-3.5, and GPT-4o. We publicly release the software pipeline that we developed to implement the proposed KGCaRe approach.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑