arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Agentic AutoRAG:通过推理驱动智能体进行RAG流水线优化

Agentic AutoRAG: RAG Pipeline Optimization through Reasoning-Driven Agents

Lasse B. Strand, Robert Jakob, Kevin O'Sullivan, Markus Kreft

arXiv 2610.08452首次发表:更新:

发表机构

ETH Zurich(苏黎世联邦理工学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出Agentic AutoRAG,一种基于LLM智能体的RAG超参数优化方法,通过检索与生成失败归因和成本权衡,在多跳QA基准上以更少试验达到更高准确率,并在医疗语料库上以更低成本实现更优性能。

AI 中文摘要

检索增强生成(RAG)是一种广泛用于将大型语言模型(LLM)锚定在外部知识中的方法。然而,配置流水线是一个昂贵的超参数优化问题,涉及许多相互作用的选项,从分块和嵌入模型到重排序和生成。现有的优化器,从贪心搜索到贝叶斯优化,将每次试验简化为一个聚合分数并进行搜索,而不对配置为何如此表现进行建模,尽管检索到的块已经提供了关于每次失败是发生在检索阶段还是检索之后阶段的证据。我们引入了Agentic AutoRAG,一个用于多目标RAG超参数优化的LLM智能体优化器,具有检索与生成失败归因功能。它提出在语料库的冻结考试上评分的配置:每次试验后,一个诊断器将每个失败的问题归因于检索或生成,而一个提议器,基于模型排名和定价的知识库,选择下一个配置,权衡准确性与成本以描绘帕累托前沿。在三个多跳问答基准上,它达到了比我们比较的所有基线更高的LLM评判准确率,并且在其前10次试验内,它匹配或超过了统计基线完整的30次试验评判准确率。在其成本感知模式下,在真实世界的医疗保健语料库上,它达到了77%的中位考试准确率,高于最强基线的71.5%,而每次查询的成本约为该基线的58%,并且它匹配该71.5%的准确率时成本约为其22%。

英文摘要

Retrieval-augmented generation (RAG) is a widely used approach for grounding large language models (LLMs) in external knowledge. However, configuring a pipeline is an expensive hyperparameter optimization problem over many interacting choices, from chunking and embedding model to reranking and generation. Existing optimizers, from greedy search to Bayesian optimization, reduce each trial to an aggregate score and search without modeling why a configuration performed as it did, even though the retrieved chunks already provide evidence about whether each failure occurred during retrieval or after it. We introduce Agentic AutoRAG, an LLM-agent optimizer for multi-objective RAG hyperparameter optimization with retrieval-versus-generation failure attribution. It proposes configurations scored on a frozen exam from the corpus: after each trial a Diagnoser attributes each failed question to retrieval or generation, and a Proposer, grounded in a knowledge base of model rankings and pricing, selects the next configuration, weighing accuracy against cost to trace a Pareto frontier. On three multi-hop QA benchmarks it reaches higher LLM-judge accuracy than every baseline we compare, and within its first 10 trials it matches or beats the statistical baselines' full 30-trial judge accuracy. In its cost-aware mode on a real-world healthcare corpus it reaches a median exam accuracy of 77%, above the strongest baseline's 71.5%, at about 58% of that baseline's cost per query, and it matches that 71.5% at about 22% of the cost.

CommentsAccepted at the Second Workshop for REsearch on Agent Language Models (REALM) at EMNLP 2026 and at the Machine Learning for Systems Workshop at NeurIPS 2026. 9 pages plus references and appendix (16 pages total), 4 figures, 6 tables. Code: https://github.com/Agentic-Systems-Lab/Agentic-AutoRAG

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑