arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.00054cs.AIcs.SE

RAG-TESTER:检索增强型大语言模型的自动化端到端测试工具

RAG-TESTER: Automated End-to-End Testing of Retrieval-Augmented Large Language Models

Ange Maiztegi, Jon Ayerdi, Miren Illarramendi, Aitor Arrieta

首次发表
浏览论文内容

中文总结 AI 辅助

RagTester是针对RAG系统的自动化端到端测试方法,通过生成测试用例、利用LLM评估结果,在24种配置中20种优于基线,可有效检测RAG的交互故障并支持部署评估。

中文摘要 AI 辅助

检索增强生成(Retrieval-Augmented Generation,RAG)使大语言模型(Large Language Models,LLMs)能够利用外部及领域特定知识,但其可靠性取决于生成模型、嵌入模型、检索机制与提示构建策略之间的交互作用。本文提出RagTester,一种针对RAG系统的自动化端到端测试方法。RagTester可生成检索文档、测试输入及预期输出,执行测试并利用LLM作为评估器对生成的答案进行评估。其测试生成策略针对复杂段落、不支持的查询及文档覆盖标准。我们使用8个LLM和6个嵌入模型对RagTester进行评估,得到24种兼容配置,并将其与基线测试输入生成器进行比较。在72000次测试执行中,RagTester检测到21633个故障,较基线多出6.6%,且在24种配置中的20种表现优于基线。检测到的故障包括检索不准确、答案不支持、检索上下文使用不完整及复杂段落解释困难等。这些结果表明,面向覆盖度的测试生成可有效暴露由检索与生成组件交互引发的故障,并支持部署前对RAG配置的评估。

英文摘要

Retrieval-Augmented Generation (RAG) enables Large Language Models (LLMs) to use external and domain-specific knowledge, but its reliability depends on the interaction between the generative model, embedding model, retrieval mechanism, and prompt construction strategy. We present RagTester, an automated end-to-end testing approach for RAG systems. RagTester generates retrieval documents, test inputs, and expected outputs; executes the tests; and evaluates the resulting answers using an LLM as a judge. Its test-generation strategy targets complex passages, unsupported queries, and document-coverage criteria. We evaluate RagTester using eight LLMs and six embedding models, yielding 24 compatible configurations, and compare it with a baseline test-input generator. Across 72,000 test executions, RagTester detected 21,633 failures, 6.6% more than the baseline, and outperformed it in 20 of the 24 configurations. The detected failures include inaccurate retrieval, unsupported answers, incomplete use of retrieved context, and difficulties interpreting complex passages. These results show that coverage-oriented test generation can effectively expose failures caused by the interaction between retrieval and generation components and support the assessment of RAG configurations before deployment.

发表机构

  • Mondragon University(蒙德拉贡大学)

机构由 AI 辅助整理,请以论文原文为准。

↑