AI 中文总结
本研究通过前瞻性审计,在固定检索策略下改变答案界面,发现重排序结论在多数设置中保持稳定,但建议在界面稳定性未验证时采用均匀平均方法。
AI 中文摘要
重排序器越来越普遍地通过下游语言模型的答案进行评估。这引发了一个检索测量问题:如果仅改变阅读器的答案界面,我们是否应该就BM25与BGE-v2-m3得出相同的结论?我们前瞻性地审计了它们在RAGuard和FEVER上对四个阅读器的声明配对效应。检索策略、证据、声明和上下文深度保持固定,而语义到标签的绑定、A/B与X/Y词汇表以及选项顺序构成了八个任务等效界面。我们探究估计的检索策略效应、其排序或选择价值是否发生变化。六个确证性设置中没有一个显示出超过预设0.015重要性阈值的统计认证界面变化,也没有一个显示出BM25-BGE排序的认证反转。在一个环境中,选择器分歧达到33.5%,但八个环境均未建立预设的重要保留价值差异。这些结果不支持广泛的复制不稳定性,但也不证明普遍不变性:五个设置仍过于不确定,无法满足预设的高阶等价条件。它们说明了为什么检索评估应区分策略层面、界面稳定性、策略排序和选择价值。当稳定性未经验证时,对枚举界面进行带有明确变异界限的均匀平均可避免偏向某一特定界面。
英文摘要
Rerankers are increasingly evaluated through downstream language-model answers. This raises a retrieval-measurement question: if only the reader's answer interface changes, should we reach the same conclusion about BM25 versus BGE-v2-m3? We prospectively audit their claim-paired effect on RAGuard and FEVER with four readers. Retrieval policies, evidence, claims, and context depth remain fixed while semantic-to-label binding, A/B versus X/Y vocabulary, and option order form eight task-equivalent interfaces. We ask whether the estimated retrieval-policy effect, its ordering, or selection value changes. None of the six confirmatory settings showed statistically certified interface variation above the prespecified 0.015 materiality threshold, and none showed a certified reversal of the BM25-BGE ordering. Selector disagreement reaches 33.5% in one environment, yet none of eight environments establishes the prespecified material held-out value difference. These results do not support broad replicated instability, but they do not prove universal invariance: five settings remain too uncertain to satisfy the prespecified higher-order equivalence condition. They show why retrieval evaluations should separate policy level, interface stability, policy ordering, and selection value. When stability is unverified, a uniform average over the enumerated interfaces with explicit variation bounds avoids privileging one interface.
Comments7 pages, 1 figure, 4 tables