arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

生成后验证主导检索优化:RAG 流水线特征的 2^4 全因子消融实验

Post-Generation Verification Dominates Retrieval Optimization: A 2^4 Factorial Ablation of RAG Pipeline Features

Ng S. T. Chong

arXiv 2609.35774首次发表:更新:

发表机构

United Nations University(联合国大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

通过 2^4 全因子消融实验发现,生成后验证(完整性检查)是 RAG 流水线中最关键的特征,其效果优于检索优化,且特征效用依赖查询类型。

AI 中文摘要

现代 RAG 流水线堆叠了许多增强特征,但这些特征通常被孤立地验证,导致它们之间的交互未被测量。我们对四个流水线特征——章节扩展(SE)、智能体搜索(AS)、完整性检查(CC)和目录引导检索(ToC)——进行了 2^4 全因子消融实验,涵盖 16 种配置、跨越八种交互类型的 24 个查询,以及两个云级模型(共 768 个条件),在五份公开文档(78-492 页)上,每个答案均对照已验证的参考进行评分。生成后验证占据主导地位:CC 是最强的特征(d=+0.48,p<0.001),同时提高了准确性、完整性和有用性,且仅 CC(4.31/5)就优于所有不含它的配置,包括三特征组合 SE+AS+ToC(4.11)。ToC 在零 LLM 成本下带来显著增益(d=+0.22);AS 效果微小且不稳定,对某些查询有帮助,对另一些则有害;SE 为中性。最高质量的配置大致使基线延迟翻倍,产生了包含六种配置的真实质量-延迟帕累托前沿。特征效用强烈依赖于查询类型——CC 在完整性要求高的查询上达到 d=+0.83——因此单一查询类型的评估会系统性地错误排序特征。我们得出结论:验证答案比优化检索更重要,且需要包含多样化查询类型的因子设计来评估 RAG 特征。

英文摘要

Modern RAG pipelines stack many enhancement features, but these features are typically validated in isolation, leaving their interactions unmeasured. We run a 2^4 full factorial ablation of four pipeline features: section expansion (SE), agentic search (AS), completeness check (CC), and table-of-contents-guided retrieval (ToC). The design crosses 16 configurations, 24 queries spanning eight interaction types, and two cloud-class models (768 scored responses) on five public documents (78-492 pages), and every answer is scored against a verified reference. On this corpus and task, post-generation verification dominates: CC is the strongest feature (d=+0.48, p<0.001), improving accuracy, completeness, and usefulness simultaneously, and CC alone (4.31/5) outperforms every configuration without it, including the three-feature SE+AS+ToC (4.11). ToC yields a significant gain at zero additional LLM calls (d=+0.22); AS is small and unstable, helping some queries and harming others; SE is neutral. The highest-quality configuration roughly doubles baseline latency, producing a genuine quality-latency Pareto frontier of six configurations. Feature utility is strongly query-type dependent (CC reaches d=+0.83 on completeness-demanding queries), so single-query-type evaluations systematically mis-rank features. We conclude that verifying answers matters more than optimizing retrieval, and that factorial designs with diverse query types are necessary to evaluate RAG features.

Comments12 pages, 11 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑