发表机构
United Nations University(联合国大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
通过 2^4 全因子消融实验发现,生成后验证(完整性检查)是 RAG 流水线中最关键的特征,其效果优于检索优化,且特征效用依赖查询类型。
AI 中文摘要
现代 RAG 流水线堆叠了许多增强特征,但这些特征通常被孤立地验证,导致它们之间的交互未被测量。我们对四个流水线特征——章节扩展(SE)、智能体搜索(AS)、完整性检查(CC)和目录引导检索(ToC)——进行了 2^4 全因子消融实验,涵盖 16 种配置、跨越八种交互类型的 24 个查询,以及两个云级模型(共 768 个条件),在五份公开文档(78-492 页)上,每个答案均对照已验证的参考进行评分。生成后验证占据主导地位:CC 是最强的特征(d=+0.48,p<0.001),同时提高了准确性、完整性和有用性,且仅 CC(4.31/5)就优于所有不含它的配置,包括三特征组合 SE+AS+ToC(4.11)。ToC 在零 LLM 成本下带来显著增益(d=+0.22);AS 效果微小且不稳定,对某些查询有帮助,对另一些则有害;SE 为中性。最高质量的配置大致使基线延迟翻倍,产生了包含六种配置的真实质量-延迟帕累托前沿。特征效用强烈依赖于查询类型——CC 在完整性要求高的查询上达到 d=+0.83——因此单一查询类型的评估会系统性地错误排序特征。我们得出结论:验证答案比优化检索更重要,且需要包含多样化查询类型的因子设计来评估 RAG 特征。
英文摘要
Modern RAG pipelines stack many enhancement features, but these features are typically validated in isolation, leaving their interactions unmeasured. We run a 2^4 full factorial ablation of four pipeline features: section expansion (SE), agentic search (AS), completeness check (CC), and table-of-contents-guided retrieval (ToC). The design crosses 16 configurations, 24 queries spanning eight interaction types, and two cloud-class models (768 scored responses) on five public documents (78-492 pages), and every answer is scored against a verified reference. On this corpus and task, post-generation verification dominates: CC is the strongest feature (d=+0.48, p<0.001), improving accuracy, completeness, and usefulness simultaneously, and CC alone (4.31/5) outperforms every configuration without it, including the three-feature SE+AS+ToC (4.11). ToC yields a significant gain at zero additional LLM calls (d=+0.22); AS is small and unstable, helping some queries and harming others; SE is neutral. The highest-quality configuration roughly doubles baseline latency, producing a genuine quality-latency Pareto frontier of six configurations. Feature utility is strongly query-type dependent (CC reaches d=+0.83 on completeness-demanding queries), so single-query-type evaluations systematically mis-rank features. We conclude that verifying answers matters more than optimizing retrieval, and that factorial designs with diverse query types are necessary to evaluate RAG features.
Comments12 pages, 11 figures