arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

一种尺寸并不适合所有情况!面向RAG的动态检索器与生成器选择

One Size Does Not Fit All! Dynamic Retriever and Generator Selection for RAG

Neeraj Anand, Payel Santra, Partha Basuchowdhuri, Debasis Ganguly, Sumit Bhatia

arXiv 2609.17709首次发表:更新:

发表机构

Media and Data Science Research Lab; Adobe Systems; IACS; University of Glasgow(媒体与数据科学研究实验室; 奥多比系统公司; 高级计算研究学院; 格拉斯哥大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对RAG固定配置效率低的问题,提出DRAG框架,通过QPP信号和微调LLM联合动态选择检索器与生成器,在多个QA基准上实现更优的效率-效果权衡。

AI 中文摘要

检索增强生成(RAG)系统通常对所有查询采用固定的检索器和生成器配置,尽管查询复杂度和信息需求存在显著差异,这导致计算资源分配效率低下。虽然检索和生成的适应性已被独立研究,但它们对端到端RAG性能的联合影响仍未得到充分探索。我们系统分析了检索器和生成器复杂度在事实性和多跳问答(QA)中的相互作用,包括桥接和组合推理任务。我们的分析表明,更强的检索通常比增加生成努力带来更大的收益,但两者都表现出收益递减和非单调性,这表明更高复杂度的配置并非对所有查询都统一更好。基于这些发现,我们提出了DRAG,一个用于选择检索器-生成器配置的查询自适应框架。我们首先提出DRAG$_{QPP}$,一种无需训练的路径选择方法,利用查询性能预测(QPP)信号指导检索器选择,并利用基于困惑度的检索上下文度量指导生成器选择。我们进一步提出DRAG$_{SFT}$,一种监督式路径选择方法,微调大语言模型(LLM)以联合预测检索器-生成器配置。在三个LLM家族和四个QA基准上,DRAG$_{QPP}$在显著降低推理延迟的同时,实现了与强静态RAG基线相当的性能,而DRAG$_{SFT}$在有效性上持续优于静态和无需训练的自适应基线。总体而言,DRAG证明了联合自适应检索和生成比静态RAG流水线实现了更优的效率-效果权衡。

英文摘要

Retrieval-Augmented Generation (RAG) systems typically employ fixed retriever and generator configurations across queries, despite substantial differences in query complexity and information needs, leading to inefficient allocation of computational resources. While retrieval and generation adaptivity have been studied independently, their joint effect on end-to-end RAG performance remains underexplored. We systematically analyze how retriever and generator complexity interacts across factoid and multi-hop question answering (QA), including bridge and composition reasoning tasks. Our analysis shows that stronger retrieval generally yields larger gains than increased generation effort, but both exhibit diminishing and non-monotonic returns, indicating that higher-complexity configurations are not uniformly better across queries. Motivated by these findings, we introduce DRAG, a query-adaptive framework for selecting retriever-generator configurations. We first propose DRAG$_\text{QPP}$, a training-free routing approach that uses Query Performance Prediction (QPP) signals to guide retriever selection and perplexity-based measures over retrieved context to guide generator selection. We further introduce DRAG$_\text{SFT}$, a supervised routing approach that fine-tunes an LLM to jointly predict retriever-generator configurations. Across three LLM families and four QA benchmarks, \qpprag~achieves performance comparable to strong static RAG baselines while substantially reducing inference latency, whereas DRAG$_\text{SFT}$ consistently improves effectiveness over static and training-free adaptive baselines. Overall, DRAG demonstrates that jointly adapting retrieval and generation achieves a more favorable effectiveness-efficiency trade-off than static RAG pipelines.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑