通过TRIAD实现多跳RAG评估自动化:从上下文提取到验证数据集生成
Automating Multi-Hop RAG Evaluation via TRIAD: From Context Extraction to Validated Dataset Generation
- University of Innsbruck(因斯布鲁克大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文提出TRIAD三阶段自动化多跳RAG评估数据集生成方法,经与MuSiQue、HotpotQA对比验证,该生成数据集可有效评估特定领域RAG系统性能,相关代码与验证结果已公开。
AI中文摘要:
大型语言模型(LLMs)的近期进展以及检索增强生成(RAG)系统在工业界的应用,催生了对特定领域问答数据集的需求,这类数据集可用于评估RAG系统在专有数据上的性能。现有数据集如HotpotQA,针对基于维基百科知识的当前RAG系统提出挑战,但无法直接迁移至特定领域场景。对RAG系统质量的全面评估需要多跳查询和无法回答的问题。本文提出TRIAD,一种三阶段自动化数据集生成方法:第一阶段为RAG系统的特定领域知识库生成问答(QA)对;第二阶段通过反馈循环对每个问答对进行验证;第三阶段为问答对扩展带有相关性标签的上下文文档,以供下游评估。我们将该方法与已建立的MuSiQue和HotpotQA数据集进行对比评估,结果显示,生成的数据集在不同RAG设置下表现出相似的性能趋势,而人工验证表明,这些问题适用于评估特定领域的RAG系统。用于生成数据集的代码及所有验证结果可在我们的GitHub仓库(https://github.com/...)获取。
英文摘要:
Recent advances in LLMs and the adoption of RAG systems in industry have created a need for domain-specific question-answer datasets that can assess RAG performance on proprietary data. Existing datasets, such as HotpotQA, challenge current RAG systems on Wikipedia-based knowledge, but they cannot be transferred directly to domain-specific settings. A comprehensive evaluation of RAG system quality requires both multi-hop queries and unanswerable questions. This paper introduces TRIAD, a three-stage automated dataset generation approach. First, it generates question--answer (QA) pairs for the domain-specific knowledge base of a RAG system. Second, a validator checks each QA-pair in a feedback loop. Third, the QA pairs are extended with relevance-labeled context documents for downstream evaluation. We evaluate this approach against the established MuSiQue and HotpotQA datasets. The results show that the generated dataset exhibits similar performance trends across different RAG setups, while human validation indicates that the questions are suitable for evaluating a domain-specific RAG system. The code used to generate the dataset and all validation results are available in our GitHub repository(https://github.com/lorenzbrehme/triad).