发表机构
Technische Hochschule Ingolstadt(英戈尔施塔特应用技术大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究通过三重鲁棒性分析,探究了不同嵌入器、语料库、评判器下GraphRAG与向量RAG的多跳可追溯性表现,发现过度引用具架构普遍性,忠实度后果受语料库影响,提出三重鲁棒性为可信RAG架构主张的最低标准。
AI 中文摘要
多项报告显示,GraphRAG在引用准确率上的表现不如向量RAG,但这种差异的具体场景与原因仍因语料库而异。我们开展了一项三重鲁棒性分析,在固定检索架构的前提下,沿三个正交维度进行变量控制:嵌入器(本地e5-small → Azure text-embedding-3-small)、语料库(DO-178C类型化边需求 → 基于MuSiQue的维基百科段落链)、评判器(配对的GPT-5.4与GPT-4.1),共完成4440次主矩阵运行、600次跨语料库运行及1200次配对忠实度判断。(C2a)过度引用具有架构普遍性:在三种设置下,GraphRAG每生成一个答案会输出11-15个ID,引用准确率为0.12-0.23,检索召回率为0.68-0.87。(C2b)其忠实度后果具有语料库条件性:在类型化边DO-178C中,GraphRAG的忠实度随跳数从74%降至40%;而在维基百科链中,同一流程的忠实度随跳数从42%升至58%,原因是过度引用的段落仍具有主题支持性。(C1)分层条件下的优胜者具有语料库条件性但嵌入器鲁棒性:在DO-178C上,普通RAG在2跳任务中表现最优;在MuSiQue上,GraphRAG在2跳任务中表现最优,两种嵌入器下结果一致。(C3)单评判器LLM的忠实度对检索状态敏感:同一评判器在不同嵌入器下的自kappa系数,GPT-5.4为0.137(41%的样本会出现判断结果变化)。仅基于稠密嵌入的学习型路由器在跳数分类上达到宏F1值0.86(C4)。我们认为,三重鲁棒性是可信RAG架构主张的最低标准。
英文摘要
GraphRAG underperforms vector RAG on citation precision in many reports, but where and why have remained corpus-bound. We present a triple-robustness analysis that holds the retrieval architecture fixed and varies three orthogonal axes embedder (local e5-small -> Azure text-embedding-3-small), corpus (DO-178C typed-edge requirements -> Wikipedia paragraph chains via MuSiQue), and judge (paired GPT-5.4 x GPT-4.1) across 4,440 main-matrix runs, 600 cross-corpus runs, and 1,200 paired faithfulness judgments. (C2a) Over-citation is architecturally universal: GraphRAG emits 11-15 IDs per answer at citation precision 0.12-0.23 and retrieval recall 0.68-0.87 across all three settings. (C2b) Its faithfulness consequence is corpus-conditional: in typed-edge DO-178C, GraphRAG faithfulness collapses 74%->40% across hops; on Wikipedia chains the same pipeline rises 42%->58% because over-cited paragraphs remain topically supporting. (C1) Stratum-conditional winners are corpus-conditional but embedder-robust: vanilla wins 2-hop on DO-178C, GraphRAG wins 2-hop on MuSiQue, identical under either embedder. (C3) Single-judge LLM faithfulness is fragile to retrieval state: same-judge self-kappa across embedders is 0.137 for GPT-5.4 (verdict change on 41% of items). A learned router on dense embeddings alone reaches macro-F1 0.86 on hop classification (C4). We argue triple-robustness is the minimum bar for trustworthy RAG architecture claims.
Comments5 pages, 3 figures, 4 tables