发表机构
University of Science and Technology of China; Metastone Technology; Information Technology Research Center, Beijing Academy of Agriculture and Forestry Sciences(中国科学技术大学; Metastone科技; 北京农林科学院信息技术研究中心)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文通过28个嵌套语料层级的受控研究,对比四类RAG范式的规模化表现,发现BM25综合性能最优,Agent+BM25全规模准确率达69.4,图基RAG存在构建瓶颈。
AI 中文摘要
检索增强生成(Retrieval-Augmented Generation,RAG)方法涵盖词汇检索、密集检索、基于图的索引及智能体搜索四类。现有研究通常在单一语料规模下的不同基准上评估这些方法,导致其准确率-成本的规模化特性尚不明确。为填补这一空白,本文对上述四种范式开展受控语料规模化研究:构建包含28个严格嵌套层级的阶梯式语料,规模从约1000篇文档扩展至512000篇,同时保持问题、固定的相关文档集及对抗文档集不变。在统一的读取器和评估协议下,我们测量官方准确率、构建与查询阶段的token用量及延迟。实验结果显示,在该受控场景中,BM25的规模化表现最佳:其在所有被测层级均处于Pareto前沿的低成本端,且从中等规模开始便主导准确率表现,无需基于大语言模型(LLM)的构建过程。文件系统智能体在最小层级上与BM25表现相当或略优,但在基准层级下每生成一个答案需多消耗39倍的查询token,且在全规模下落后近20个百分点。匹配检索的替换方案可逆转该缺陷:在相同150个问题上,Agent+BM25在全规模下得分为69.4,而原始文件智能体得分为36.9,原生BM25得分为54.8。基于图的RAG则遭遇构建瓶颈:其最重的构建器每索引1个语料token最多消耗24.6个生成式LLM token,却仅能覆盖全语料的前2%,而可扩展变体在相同层级下的准确率仍低于BM25。
英文摘要
Retrieval-augmented generation (RAG) spans lexical and dense retrieval, graph-based indexing, and agentic search, but these paradigms are usually evaluated on different benchmarks at one corpus size, leaving their accuracy-cost scaling unclear. To bridge this gap, we present a controlled study that varies corpus size along 28 strictly nested tiers spanning roughly 450-fold, while holding questions and a fixed bedrock of relevant and adversarial documents unchanged. Under one reader model and one judging protocol, we measure official accuracy, construction and query tokens, and latency. The results reveal a scale-dependent crossover rather than an unconditional winner. File-System Agent leads at the smallest shared tiers, but its sequential exploration costs 39 times more query tokens at the bedrock and becomes less effective as the search space grows. Around 10 million corpus tokens, BM25 overtakes it and leads at every larger shared tier, with a margin approaching 20 points at full scale. BM25 also anchors the low-cost end of the Pareto frontier without LLM-based construction. Dense retrieval remains efficient but less accurate, whereas graph-based RAG encounters construction walls before deployment scale and its scalable variants remain below BM25 at shared tiers. Overall, corpus growth increasingly favors global candidate ranking: lexical retrieval is the strongest scalable default, while agentic reasoning works best after ranked discovery rather than in place of it.