arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.25115cs.CLcs.IR

少即是多:通过证据前置和压力自适应预算缓解RAG瓶颈

Less can be More: Relieving RAG Bottlenecks via Evidence Frontloading and Pressure-Adaptive Budgeting

Weibin Cai, Reza Zafarani

首次发表
浏览论文内容

中文总结 AI 辅助

该研究针对RAG系统瓶颈转移问题,提出无训练框架PACE,结合证据前置与压力自适应预算,在多跳问答数据集及模拟中提升了证据召回率并降低了重排序密集型负载的延迟。

中文摘要 AI 辅助

现有提升检索增强生成(RAG)效率的方法主要优化下游大语言模型(LLM)生成,如上下文压缩或服务优化。但RAG是端到端系统,其瓶颈会随不同服务负载在上游重排序和下游生成间转移,当查询率高或重排序预算大时,上游重排序会成为主导瓶颈。减少重排序预算可缓解该瓶颈,但可能丢失支撑证据、降低召回率。为解决此问题,本文提出PACE(Prioritized Adaptive Coverage of Evidence,证据的优先自适应覆盖),这是一种无训练框架,结合证据前置与压力自适应预算。PACE先通过边际证据覆盖度对候选文档重新排序,优先选择与查询相关、互补且对形成多跳证据链有用的文档,该目标是单调次模函数,使贪心选择有(1-1/e)近似保证。PACE再根据重排序器与LLM的相对压力动态调整重排序预算。在3个多跳问答数据集及在线服务模拟上的实验显示,PACE可提升证据召回率、降低重排序密集型工作负载下的p95延迟,更重要的是,两个组件共同揭示了“少即是多”:证据密集的排名靠前候选文档,用更少重排序文档即可实现更高最终召回率。

英文摘要

Existing methods for improving Retrieval-Augmented Generation (RAG) efficiency mainly optimize downstream LLM generation, such as context compression or serving optimization. However, RAG is an end-to-end system, and its bottleneck can shift between upstream reranking and downstream generation under different serving loads and reranking budgets.In this paper, we first empirically characterize this shifting-bottleneck behavior and show that upstream reranking can become the dominant bottleneck under high query rates or large reranking budgets. Reducing the reranking budget can relieve this bottleneck, but it may also drop supporting evidence and degrade recall. To address this problem, we propose \textbf{\textsf{PACE}} (\textbf{P}rioritized \textbf{A}daptive \textbf{C}overage of \textbf{E}vidence), a training-free framework that combines \textit{evidence frontloading} with \textit{pressure-adaptive budgeting}. \textsf{PACE} first reorders candidates by marginal evidence coverage, prioritizing documents that are query-relevant, complementary, and useful for forming multi-hop evidence chains. We show that this objective is monotone submodular, giving greedy selection a $(1-1/e)$ approximation guarantee. \textsf{PACE} then dynamically adjusts the reranking budget according to the relative pressure of the reranker and the LLM. Experiments on three multi-hop QA datasets and online serving simulations show that \textsf{PACE} improves evidence recall, reduces p95 latency under ranking-heavy workloads. More importantly, the two components together reveal that \textit{less can be more}: an evidence-dense top-ranked candidates enable higher final recall with fewer reranked documents.

发表机构

  • Syracuse University(雪城大学)

机构由 AI 辅助整理,请以论文原文为准。

↑