检索增强生成的不可回答性评估
Unanswerability Evaluation for Retrieval Augmented Generation
- Salesforce Research(Salesforce 研究)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文提出UAEval4RAG框架,用于评估RAG系统对不可回答查询的处理能力,通过六类分类和自动合成查询,揭示组件选择与提示设计对平衡准确性与拒绝率的关键作用。
AI中文摘要:
现有的检索增强生成(RAG)系统评估框架主要关注可回答的查询,但忽视了适当拒绝不可回答请求的重要性。本文中,我们引入了UAEval4RAG,一个旨在评估RAG系统能否有效处理不可回答查询的框架。我们定义了一个包含六种不可回答类别的分类体系,UAEval4RAG能够针对任意给定的知识库自动合成多样且具有挑战性的查询,并采用未回答率和可接受率指标进行衡量。我们使用各种RAG组件进行了实验,包括检索模型、重写方法、重排序器、语言模型和提示策略,揭示了RAG系统性能中隐藏的权衡。我们的发现强调了组件选择和提示设计在优化RAG系统中的关键作用,以平衡可回答查询的准确性与不可回答查询的高拒绝率。UAEval4RAG为开发更健壮、更可靠的RAG系统提供了宝贵的见解和工具。
英文摘要:
Existing evaluation frameworks for retrieval-augmented generation (RAG) systems focus on answerable queries, but they overlook the importance of appropriately rejecting unanswerable requests. In this paper, we introduce UAEval4RAG, a framework designed to evaluate whether RAG systems can handle unanswerable queries effectively. We define a taxonomy with six unanswerable categories, and UAEval4RAG automatically synthesizes diverse and challenging queries for any given knowledge base with unanswered ratio and acceptable ratio metrics. We conduct experiments with various RAG components, including retrieval models, rewriting methods, rerankers, language models, and prompting strategies, and reveal hidden trade-offs in performance of RAG systems. Our findings highlight the critical role of component selection and prompt design in optimizing RAG systems to balance the accuracy of answerable queries with high rejection rates of unanswerable ones. UAEval4RAG provides valuable insights and tools for developing more robust and reliable RAG systems.