arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

当知识发生变化时:面向RAG系统的变异体态测试

When Knowledge Changes: Metamorphic Testing of RAG Systems with Mutations

Jinhan Kim, Samuele Pasini, Paolo Tonella

arXiv 2607.26843首次发表:更新:

AI 中文总结

针对现有RAG系统评估方法无法检测语料库演变带来的故障问题,提出含11个变异算子的体态测试框架,实验显示其F1分数远优于RAGAS,可有效评估RAG系统在语料库变化下的一致性。

AI 中文摘要

基于检索增强生成(RAG)的大语言模型(LLM)系统依赖于会随时间演变变化的外部文档语料库。然而,当前的评估方法(如RAGAS)仅针对静态快照评估正确性,无法检测常规更新、事实变更或噪声改变底层数据时产生的故障。我们提出一种用于评估RAG系统在语料库演变下一致性的变异体态测试框架,形式化了故障分类体系与11个变异算子,这些算子会在分块前(检索索引)和分块后(检索上下文)层面对系统进行系统性扰动。针对5个数据集及超过2.8万个变异体的实证评估显示,变异违反率为4.9%-10.2%。在针对真实值的元评估中,我们的变异神谕(metamorphic oracle)的F1分数达0.927-1.000,而最优的RAGAS指标仅为0.570。最后,我们就通过检索重新配置、生成器升级及基于LLM的重排序缓解这些故障提供了可行见解。

英文摘要

Retrieval-Augmented Generation (RAG)-based LLM systems rely on external document corpora that can evolve and change over time. However, current evaluation methodologies (e.g., RAGAS) assess correctness against static snapshots, failing to detect faults when routine updates, factual changes, or noise alter the underlying data. We introduce a metamorphic testing framework that evaluates the consistency of RAG systems under corpus evolution. We formalise a fault taxonomy and 11 mutation operators that systematically perturb the system at both the pre-chunk (retrieval index) and post-chunk (retrieved context) levels. An empirical evaluation across five datasets and over 28k mutants reveals metamorphic violation rates of 4.9-10.2%. In a meta-evaluation against ground truth, our metamorphic oracle achieves F1 scores of 0.927-1.000, while the best RAGAS metric reaches only 0.570. Finally, we provide actionable insights into mitigating these faults through retrieval re-configuration, generator upgrades, and LLM-based reranking.

CommentsASE 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑