发表机构
Yale School of Medicine(耶鲁医学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出RAGFlip,研究检索器升级中查询级负面翻转现象,发现替代检索器虽提升整体覆盖率但导致部分BM25成功查询回退,建议评估时结合查询级兼容性。
AI 中文摘要
检索器升级通常使用聚合指标进行评估,这可能会掩盖先前检索器已正确服务的查询上的性能回退。我们将这些回退研究为负面翻转:即BM25检索到判定为相关的段落而替代检索器未检索到的查询。我们在三个BEIR数据集(Natural Questions、HotpotQA和FiQA)上评估了BGE-large、E5-large-v2和SPLADE,跨越五个检索深度。所有替代检索器都提高了整体检索覆盖率。负面翻转在每个设置中都会发生,并且因语料库、检索器和深度而有很大差异。在k=1时,在所评估的设置中,8.6%-37.5%的BM25成功案例丢失。负面翻转率在较大深度设置中较低,其中BM25支持的队列在每个深度分别定义。这些比率使用任意相关支持标签。在HotpotQA上,在k=10时,要求每个正向qrel段落将负面翻转率提高到12.3%-17.3%。与BM25的简单固定预算组合减少了这些回退,而HotpotQA阅读器实验提供了一个有限的下游检查,其中一些检索翻转伴随着答案回退。这些结果促使在评估检索器更新时,将查询级兼容性与聚合检索质量一起考虑。
英文摘要
Retriever upgrades are typically evaluated using aggregate metrics, which can hide regressions on queries the previous retriever already served correctly. We study these regressions as negative flips: queries for which BM25 retrieves a judged relevant passage and the replacement does not. We evaluate BGE-large, E5-large-v2, and SPLADE on three BEIR collections: Natural Questions, HotpotQA, and FiQA, across five retrieval depths. All replacements improve overall retrieval coverage. Negative flips occur in every setting and vary substantially by corpus, retriever, and depth. At k=1, 8.6-37.5% of BM25 successes are lost across the evaluated settings. Negative-flip rates are lower in the larger-depth settings, where the BM25-supported cohort is defined separately at each depth. These rates use the any-relevant support label. On HotpotQA at k=10, requiring every positive qrel passage raises the negative-flip rate to 12.3-17.3%. Simple fixed-budget combinations with BM25 reduce these regressions, and a HotpotQA reader experiment provides a limited downstream check in which some retrieval flips are accompanied by answer regressions. These results motivate evaluating retriever updates using query-level compatibility alongside aggregate retrieval quality.
Comments19 pages, 4 figures. Code available at https://github.com/Elyasirankhah/RAGFlip