发表机构
Turkish Aerospace Industries; Technical University of Munich(土耳其航空航天工业公司; 慕尼黑工业大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文针对多跳需求可追溯性,开展三重鲁棒性分析,对比不同RAG架构、嵌入器等的表现,发现现有结论分歧的原因,提出需按该鲁棒性标准验证RAG架构主张。
AI 中文摘要
已有的GraphRAG与向量RAG的对比结论存在分歧,且相关证据通常仅关联单一语料库、嵌入器和评判器,我们还发现其与引文质量的测量位置有关。本文提出一种三重鲁棒性分析方法,固定五管道架构矩阵,改变嵌入器(本地e5-small vs. Azure text-embedding-3-small)、语料库(DO-178C类型边需求 vs. 基于MuSiQue的维基百科段落链)和评判器(两个语料库上配对的GPT-5.4与GPT-4.1),开展2×4440次主矩阵运行、600次跨语料库运行及5000余次忠实性评判。结果显示:(C2a)GraphRAG的图遍历在精度0.12-0.23时会占满上下文窗口,但合成器在精度0.48-0.65时会选择性引用;将检索集作为归因集评分会反转架构排名,这调和了部分现有研究的分歧。(C1)答案级引文优胜者受语料库和分层条件影响,但对嵌入器具有鲁棒性:GraphRAG在短跳DO-178C查询上与基准表现相当,在所有MuSiQue分层中胜出,而智能体管道仅在3跳及以上的需求查询中表现领先。(C2b)忠实性受语料库条件影响:在DO-178C上,忠实性随跳数距离下降(4种评判器×嵌入器组合中有3种的趋势p<0.05);在维基百科链上,两种评判器均未出现性能崩溃。(C3)单评判器LLM的忠实性对检索状态敏感:GPT-5.4在不同嵌入器间的自身kappa值为0.137(41%的结论变化),而同一天的重测基准为0.76;11周后对冻结输入重新评判,两种评判器的kappa值均≤0.14。(C4)仅基于稠密嵌入的学习路由器在跳数分类上达到宏F1值0.86。本文认为,RAG架构的相关结论需在该鲁棒性水平(包括对引文测量点的鲁棒性)下经过测试,才可被信任。
英文摘要
Reported verdicts on GraphRAG versus vector RAG disagree, and the evidence is typically tied to a single corpus, embedder, and judge -- and, we show, to where citation quality is measured. We present a triple-robustness analysis that holds a five-pipeline architecture matrix fixed and varies embedder (local e5-small vs. Azure text-embedding-3-small), corpus (DO-178C typed-edge requirements vs. Wikipedia paragraph chains via MuSiQue), and judge (paired GPT-5.4 x GPT-4.1 on both corpora), over 2x4,440 main-matrix runs, 600 cross-corpus runs, and over 5,000 faithfulness judgments. (C2a) GraphRAG's graph walk floods the context window at precision 0.12-0.23, but the synthesizer cites selectively at precision 0.48-0.65; scoring the retrieved set as the attribution set inverts the architecture ranking, which reconciles part of the disagreement in prior reports. (C1) Answer-level citation winners are corpus- and stratum-conditional but embedder-robust: GraphRAG ties vanilla on short-hop DO-178C queries and wins every MuSiQue stratum, while agentic pipelines lead only on 3+-hop requirements queries. (C2b) Faithfulness is corpus-conditional: on DO-178C it declines with hop distance (trend p<0.05 in three of four judge x embedder combinations); on Wikipedia chains neither judge shows a collapse. (C3) Single-judge LLM faithfulness is fragile to retrieval state: GPT-5.4's self-kappa across embedders is 0.137 (41% verdict change) against a same-day test-retest floor of 0.76, and re-judging frozen inputs eleven weeks later gives kappa <= 0.14 for both judges. A learned router on dense embeddings alone reaches macro-F1 0.86 on hop classification (C4). We argue that RAG architecture claims should be tested at this level of robustness -- including robustness to the citation-measurement point -- before they are trusted.
Comments6 pages, 3 figures, 4 tables