图式推理瓶颈:用于多跳问答的结构化提示与上下文压缩
The Reasoning Bottleneck in Graph-RAG: Structured Prompting and Context Compression for Multi-Hop QA
- University of Victoria(维多利亚大学)
- Santa Clara University(圣克拉拉大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
图式推理瓶颈:本文提出结构化提示和上下文压缩方法,通过SPARQL链式思考提示和图遍历压缩提升多跳问答准确性,验证其在不同Graph-RAG系统中的有效性。
AI中文摘要:
图式推理系统通过将文档索引到知识图中实现强大的多跳问答,但强大的检索不保证强大的答案。在三个多跳问答基准(HotpotQA、MuSiQue、2WikiMultiHopQA)上评估KET-RAG系统,发现77%至91%的问题答案在检索上下文中,但准确性仅为35%至78%,且73%至84%的错误是推理失败。我们提出两种增强方法:(i) SPARQL链式思考提示,将问题分解为与实体关系上下文对齐的三元组查询;(ii) 图遍历压缩,通过知识图遍历将上下文压缩约60%而无需LLM调用。SPARQL CoT将准确性提高2至14个百分点;图遍历压缩在配对结构化提示时在较小模型上平均增加6个百分点。令人惊讶的是,我们证明通过问题类型路由,一个完全增强的预算开放权重Llama-8B模型在所有三个基准上以约12倍更低的成本匹配或超过未增强的Llama-70B基线。在LightRAG上的复制实验验证了我们的增强方法在不同Graph-RAG系统中的泛化能力。
英文摘要:
Graph-RAG systems achieve strong multi-hop question answering by indexing documents into knowledge graphs, but strong retrieval does not guarantee strong answers. Evaluating KET-RAG, a leading Graph-RAG system, on three multi-hop QA benchmarks (HotpotQA, MuSiQue, 2WikiMultiHopQA), we find that 77% to 91% of questions have the gold answer in the retrieved context, yet accuracy is only 35% to 78%, and 73% to 84% of errors are reasoning failures. We propose two augmentations: (i) SPARQL chain-of-thought prompting, which decomposes questions into triple-pattern queries aligned with the entity-relationship context, and (ii) graph-walk compression, which compresses the context by ~60% via knowledge-graph traversal with no LLM calls. SPARQL CoT improves accuracy by +2 to +14 pp; graph-walk compression adds +6 pp on average when paired with structured prompting on smaller models. Surprisingly, we show that, with question-type routing, a fully augmented budget open-weight Llama-8B model matches or exceeds the unaugmented Llama-70B baseline on all three benchmarks at ~12x lower cost. A replication on LightRAG confirms that our augmentations transfer across Graph-RAG systems.