NeSy-RAG:用于可解释问答的神经符号检索增强生成框架
NeSy-RAG: Neuro-Symbolic RAG for Explainable Question Answering
- Heidelberg University(海德堡大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
NeSy-RAG是一种模块化神经符号RAG框架,可生成透明推理轨迹,在ShARC基准上以61.1%的准确率优于同模型RAG基线,实现可解释问答。
AI中文摘要:
检索增强生成(RAG)通过将大型语言模型(LLM)与文本语料库等外部知识绑定,改进了问答任务,但它的推理过程仍在很大程度上不透明:中间推理步骤难以验证,且无法可靠地归因于特定证据。此外,用户特定上下文的缺失很少被系统检测到,往往导致输出不完整或不正确。我们提出NeSy-RAG,一种模块化神经符号RAG框架,它从检索到的文本块中合成可归因的Prolog模块。对于每个文本块,系统生成语义上有意义的谓词,这些谓词编码布尔声明,可能依赖于用户事实。利用联合自然语言-代码嵌入,谓词被检索并组合成Prolog查询。为解决用户上下文不完整的问题,我们引入了一种符号知识缺口检测机制,该机制识别其真值会影响查询结果的缺失用户事实,并自动触发后续交互。执行生成的Prolog查询会产生确定性答案以及透明的执行轨迹,该轨迹将每个推理步骤与其原始来源关联起来。在ShARC基准上,无需针对特定领域进行训练,NeSy-RAG达到了61.1%的准确率,优于达到42.8%准确率的同模型RAG基线。
英文摘要:
Retrieval-augmented generation (RAG) improves question answering by grounding large language models (LLMs) in external knowledge such as text corpora. However, its reasoning process remains largely opaque: intermediate reasoning steps are difficult to verify and cannot be reliably attributed to specific evidence. Moreover, missing user-specific context is rarely detected systematically, often leading to incomplete or incorrect output. We propose NeSy-RAG, a modular neuro-symbolic RAG framework that synthesizes attributable Prolog modules from retrieved text chunks. For each chunk, the system generates semantically meaningful predicates that encode Boolean claims, which may depend on user facts. Using joint natural language-code embeddings, predicates are retrieved and composed into Prolog queries. To address incomplete user context, we introduce a symbolic knowledge-gap detection mechanism that identifies missing user facts whose truth values affect the query outcome and automatically triggers follow-up interactions. Executing the resulting Prolog queries yields deterministic answers together with transparent execution traces that link each reasoning step to its originating source. On the ShARC benchmark, without domain-specific training, NeSy-RAG achieves 61.1% accuracy, outperforming a same-model RAG baseline that achieves 42.8% accuracy.