叠加提示:改进并加速检索增强生成
Superposition Prompting: Improving and Accelerating Retrieval-Augmented Generation
- Apple(苹果公司)
- Meta
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对长上下文推理成本高和无关上下文干扰问题,提出叠加提示方法,无需微调即可并行处理文档路径,在多个问答基准上同时提升时间效率与准确率,如计算时间减少93倍且准确率提升43%。
AI中文摘要:
尽管大型语言模型(LLM)取得了成功,但它们也表现出显著的缺点,尤其是在处理长上下文时。其推理成本随序列长度呈二次方增长,使得在诸如检索增强生成(RAG)等某些实际文本处理应用中部署成本高昂。此外,LLM还表现出“分心现象”,即提示中的无关上下文会降低输出质量。为了解决这些缺点,我们提出了一种新颖的RAG提示方法——叠加提示(superposition prompting),该方法可直接应用于预训练的基于Transformer的LLM,而无需微调。在高层次上,叠加提示允许LLM并行处理输入文档的提示路径,一旦路径被判定为无关则将其丢弃。我们证明了我们的方法能够在多个预训练LLM上同时提高各种问答基准的时间效率。此外,当检索到的上下文相对于模型训练时的上下文较大时,我们的技术显著提高了准确性。例如,在NaturalQuestions-Open数据集上,使用MPT-7B指令微调模型,我们的方法相较于朴素RAG,实现了93倍的计算时间减少,同时准确率提高了43%。
英文摘要:
Despite the successes of large language models (LLMs), they exhibit significant drawbacks, particularly when processing long contexts. Their inference cost scales quadratically with respect to sequence length, making it expensive for deployment in some real-world text processing applications, such as retrieval-augmented generation (RAG). Additionally, LLMs also exhibit the "distraction phenomenon", where irrelevant context in the prompt degrades output quality. To address these drawbacks, we propose a novel RAG prompting methodology, *superposition prompting*, which can be directly applied to pre-trained transformer-based LLMs *without the need for fine-tuning*. At a high level, superposition prompting allows the LLM to process input documents in parallel *prompt paths*, discarding paths once they are deemed irrelevant. We demonstrate the capability of our method to simultaneously enhance time efficiency across a variety of question-answering benchmarks using multiple pre-trained LLMs. Furthermore, our technique significantly improves accuracy when the retrieved context is large relative the context the model was trained on. For example, our approach facilitates a 93x reduction in compute time while *improving* accuracy by 43% on the NaturalQuestions-Open dataset with the MPT-7B instruction-tuned model over naive RAG.