通过稀疏上下文选择加速检索增强生成的推理
Accelerating Inference of Retrieval-Augmented Generation via Sparse Context Selection
- Google DeepMind(谷歌DeepMind)
- University of California, Los Angeles(加利福尼亚大学洛杉矶分校)
- Google(谷歌)
- Université de Montréal(蒙特利尔大学)
- Mila(米拉研究所)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文提出Sparse RAG,通过并行编码检索文档并仅关注相关缓存,减少解码加载量,加速推理并提升生成质量,在两类任务上实现效率与质量的最佳平衡。
AI中文摘要:
大型语言模型(LLMs)通过引入外部上下文进行检索增强,展现出稳健的性能和广泛的适用性。然而,输入长度随检索文档数量线性增长,导致延迟显著增加。本文提出一种名为Sparse RAG的新范式,旨在通过稀疏性降低计算成本。具体而言,Sparse RAG并行编码检索到的文档,从而消除了由检索文档的长距离注意力带来的延迟。随后,LLMs仅通过自回归方式关注高度相关的缓存来选择性地解码输出,这些缓存是通过使用特殊控制标记提示LLMs来选择的。值得注意的是,Sparse RAG将每个单独文档的评估与响应的生成合并到单一过程中。RAG系统中设计的稀疏机制有助于减少解码期间加载的文档数量,从而加速RAG系统的推理。此外,过滤掉不相关的上下文增强了模型对相关上下文的关注,从本质上提高了其生成质量。在两个数据集上的评估结果表明,Sparse RAG能够在生成质量与计算效率之间取得最佳平衡,并展示了其在短文本和长文本生成任务中的泛化能力。
英文摘要:
Large language models (LLMs) augmented with retrieval exhibit robust performance and extensive versatility by incorporating external contexts. However, the input length grows linearly in the number of retrieved documents, causing a dramatic increase in latency. In this paper, we propose a novel paradigm named Sparse RAG, which seeks to cut computation costs through sparsity. Specifically, Sparse RAG encodes retrieved documents in parallel, which eliminates latency introduced by long-range attention of retrieved documents. Then, LLMs selectively decode the output by only attending to highly relevant caches auto-regressively, which are chosen via prompting LLMs with special control tokens. It is notable that Sparse RAG combines the assessment of each individual document and the generation of the response into a single process. The designed sparse mechanism in a RAG system can facilitate the reduction of the number of documents loaded during decoding for accelerating the inference of the RAG system. Additionally, filtering out undesirable contexts enhances the model's focus on relevant context, inherently improving its generation quality. Evaluation results of two datasets show that Sparse RAG can strike an optimal balance between generation quality and computational efficiency, demonstrating its generalizability across both short- and long-form generation tasks.