RAG-Safety-Bench:检索增强大语言模型安全性的可靠评估
RAG-Safety-Bench: Reliable Evaluation of Retrieval-Augmented LLM Safety
浏览论文内容
中文总结 AI 辅助
提出RAG-Safety-Bench基准,通过四种条件隔离检索质量影响,评估RAG对LLM安全性的影响,发现良性能力与不安全能力呈反比,基线安全护栏无法保证RAG下游安全。
中文摘要 AI 辅助
允许大语言模型(LLM)从一组可信文档中检索信息可以提高可靠性并减少幻觉。然而,最近的研究表明,当被提示有害或危险内容时,检索增强生成(RAG)可能对生成响应的整体安全性产生意想不到的副作用。随着越来越多的最终用户转向RAG,将企业文档和知识库整合到基于LLM的系统中,需要更清晰地理解导致这一结果的机制。我们引入了RAG-Safety-Bench,一个用于衡量RAG对LLM模型安全性影响的基准。通过消除检索器质量的混淆效应,并将问题清晰地划分为四种条件——非RAG、使用包含有害请求答案的oracle文档的RAG、使用与有害请求相关但不包含具体答案的文档的RAG,以及使用随机安全文档的RAG——该基准隔离了观察到的安全性退化中不同因素的影响。我们报告了五个开源LLM的结果,显示出良性能力与不安全能力之间的反比关系,强有力证据表明基线安全护栏在RAG情况下并不能带来下游安全性保证,以及模型特定的证据支持先前发现,即即使是良性文档也可能在启用检索的系统中导致不安全生成。
英文摘要
Allowing large language models (LLMs) to retrieve information from a set of trusted documents can increase reliability and reduce hallucination. However, recent work has demonstrated that retrieval-augmented generation (RAG) can have unintended side effects on the overall safety of the generated responses, when prompted for harmful or dangerous content. A clearer understanding of the mechanisms leading to this result is needed, as increasing numbers of end users turn to RAG to incorporate corporate documents and knowledge bases into LLM-based systems. We introduce RAG-Safety-Bench, a benchmark to measure the safety impact of RAG on LLM models. By removing the confounding effect of retriever quality, and cleanly separating the problem into four conditions -- non-RAG, RAG with an oracle document containing the answer to the harmful request, RAG with documents related to the harmful request but without the specific answer, and RAG with random, safe documents -- the benchmark isolates the impacts of different factors in the observed safety degradation. We report results across five open-source LLMs, showing an inverse relationship between benign and unsafe capability, strong evidence that baseline safety guardrails do not lead to downstream safety guarantees in the RAG case, and model-specific support for previous findings that even benign documents can lead to unsafe generation in retrieval-enabled systems.
发表机构
- University of Ottawa(渥太华大学)
机构由 AI 辅助整理,请以论文原文为准。