基于检索增强生成的大语言模型中的提示扰动研究
Prompt Perturbation in Retrieval-Augmented Generation based Large Language Models
- The University of New South Wales(新南威尔士大学)
- CSIRO Data61(澳大利亚联邦科学与工业研究组织数据61)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本研究发现短前缀提示扰动会大幅降低RAG大模型输出准确性,提出GGPP优化技术可高效诱导错误输出,同时基于神经元激活差异训练检测器提升模型鲁棒性。
AI中文摘要:
随着大语言模型(LLMs)在各领域的应用快速增长,其鲁棒性变得愈发重要。检索增强生成(Retrieval-Augmented Generation,RAG)被视为提升大语言模型文本生成可信度的一种手段。然而,输入的微小差异会如何影响基于RAG的大语言模型的输出,目前尚未得到充分研究。\n在本研究中,我们发现即使在提示中插入一个很短的前缀,也会导致生成的输出与事实正确答案相差甚远。我们通过引入一种名为梯度引导提示扰动(Gradient Guided Prompt Perturbation,GGPP)的新型优化技术,系统评估了此类前缀对RAG的影响。GGPP在引导基于RAG的大语言模型输出转向目标错误答案方面实现了很高的成功率,它还能应对提示中要求忽略无关上下文的指令。\n我们还利用大语言模型在有无GGPP扰动的提示之间的神经元激活差异,提出了一种提升基于RAG的大语言模型鲁棒性的方法:通过在GGPP生成的提示所触发的神经元激活上训练一个高效的检测器来实现。我们在开源大语言模型上的评估证明了所提方法的有效性。
英文摘要:
The robustness of large language models (LLMs) becomes increasingly important as their use rapidly grows in a wide range of domains. Retrieval-Augmented Generation (RAG) is considered as a means to improve the trustworthiness of text generation from LLMs. However, how the outputs from RAG-based LLMs are affected by slightly different inputs is not well studied. In this work, we find that the insertion of even a short prefix to the prompt leads to the generation of outputs far away from factually correct answers. We systematically evaluate the effect of such prefixes on RAG by introducing a novel optimization technique called Gradient Guided Prompt Perturbation (GGPP). GGPP achieves a high success rate in steering outputs of RAG-based LLMs to targeted wrong answers. It can also cope with instructions in the prompts requesting to ignore irrelevant context. We also exploit LLMs' neuron activation difference between prompts with and without GGPP perturbations to give a method that improves the robustness of RAG-based LLMs through a highly effective detector trained on neuron activation triggered by GGPP generated prompts. Our evaluation on open-sourced LLMs demonstrates the effectiveness of our methods.