arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2605.17989cs.CLcs.AI

检索增强生成的预测预取

Predictive Prefetching for Retrieval-Augmented Generation

  • Department of Computer Science, University of Massachusetts Boston(马萨诸塞大学波士顿分校计算机科学系)

机构由 AI 辅助整理,请以论文原文为准。

Wuyang Zhang, Shichao Pei

更新

AI总结:

本文提出了一种先进的异步检索框架,通过预测检索触发时机和所需信息,以减少延迟并提高生成效率,同时保持回答质量。

AI中文摘要:

检索增强生成(RAG)通过在大型语言模型中增强事实性,但因其同步检索导致显著延迟。尽管近期工作探索了异步检索,但现有方法依赖于检索与生成之间的启发式协调,并假设解码期间信息需求稳定,这在复杂、多领域设置中往往失效。本文提出了一种先进的异步检索框架,该框架能够与不断演变的信息需求相匹配,通过利用生成动态中出现的语义前驱,使用三个组件——检索预测器、上下文监视器和查询生成器,显式预测何时应触发检索以及应检索什么信息。在多个基准测试上的实验表明,该方法可实现高达43.5%的端到端延迟减少和62.4%的时间到第一个token的提升,同时保持与同步RAG基线相当的回答质量。

英文摘要:

Retrieval-Augmented Generation (RAG) improves factual grounding in large language models but suffers from substantial latency due to synchronous retrieval. While recent work explores asynchronous retrieval, existing approaches rely on heuristic coordination between retrieval and generation and assume stable information demands during decoding that often break in complex, multi-domain settings. In this paper, we propose an advanced asynchronous retrieval framework that enables predictive prefetching aligned with evolving information needs. The framework explicitly predicts when retrieval should be triggered and what information should be retrieved using three components, a retrieval predictor, a context monitor, and a query generator, by exploiting semantic precursors in generation dynamics that emerge several tokens before uncertainty becomes critical. Experiments on multiple benchmarks demonstrate up to 43.5% end-to-end latency reduction and 62.4% improvement in time-to-first-token, while maintaining answer quality comparable to synchronous RAG baselines.

补充信息

↑