AI 中文总结
研究用块覆盖率(CC)测试检索增强生成(RAG)系统的检索组件,CC衡量语料块检索比例,可指导测试选择与生成,通过实验表明CC引导测试能更快达覆盖率且提高故障检测有效性,无需测试预言机就能捕捉检索多样性。
AI 中文摘要
基于检索增强生成(RAG)的系统越来越多地部署在高风险环境中,其正确行为不仅取决于语言模型,还取决于推理时选择外部文档的检索组件。现有RAG评估指标按查询评估检索和生成质量,对测试套件是否充分检验系统整体检索行为洞察有限。本文引入块覆盖率(CC),一种独立于预言机的测试充分性标准来测试RAG系统的检索组件。CC衡量测试套件中至少被检索一次的语料块比例,提供已被检验的检索空间部分的结构视图。还展示了CC如何通过优先选择扩展未被检验检索区域覆盖范围的查询来指导测试选择和生成。在临床和金融RAG系统场景中评估CC,CC引导的测试比随机选择快1.7倍、比冗余偏向策略快4.2倍达到可实现覆盖率的50%,且将故障检测有效性(APFD)比随机提高10%到25%。结果表明CC无需测试预言机就能捕捉与有效测试相关的检索多样性。
英文摘要
Retrieval-Augmented Generation (RAG)-based systems\footnote{For brevity, RAG-based systems are referred to as RAG systems throughout this paper.} are increasingly deployed in high-stakes settings where correct behaviour depends not only on the language model but also on the retrieval component that selects external documents at inference time. While existing RAG evaluation metrics assess retrieval and generation quality on a per-query basis, typically relying on query-level test oracles such as reference answers or relevance annotations, they provide limited insight into whether a test suite adequately exercises the retrieval behaviour of the system as a whole. In this paper, we introduce Chunk Coverage (CC), an oracle-independent test adequacy criterion for testing the retrieval component of RAG systems. CC measures the fraction of corpus chunks that are retrieved at least once across a test suite, providing a structural view of which parts of the retrieval space have been exercised. We further show how CC can be used to guide test selection and generation by prioritising queries that expand coverage of previously unexercised retrieval regions. We evaluate CC on clinical and financial RAG system scenarios. CC-guided testing reaches 50% of attainable coverage 1.7x faster than random selection and 4.2x faster than redundancy-biased strategies. Moreover, CC improves fault detection effectiveness (APFD) by 10% to 25% over random, indicating earlier discovery of distinct retrieval faults. These results show that CC captures retrieval diversity relevant to effective testing without requiring test oracles.
CommentsISSTA 2026