幅度幻象:重新思考面向推理密集型检索的置信度
The Magnitude Mirage: Rethinking Confidence for Reasoning-Intensive Retrieval
浏览论文内容
中文总结 AI 辅助
针对RAG检索弃权(不执行)中幅度置信度的失效问题,提出采用分数分布信号(如分数差距和LSMV)替代幅度阈值,在逻辑与时间推理任务上显著提升弃权(不执行)AUROC,且零额外成本。
中文摘要 AI 辅助
许多生产级RAG系统通过设定原始相似度分数的阈值来实现检索弃权(不执行),隐含地将分数幅度视为置信度信号。我们证明,当查询需要超越语义匹配的推理时,这种做法会系统性退化。在11种检索架构和28个数据集上,神经检索器始终对语义相关但违反约束的文档赋予高相似度分数,导致基于幅度的阈值在逻辑和时间推理任务上崩溃至接近随机的弃权(不执行)性能——我们将这一失败称为“幅度幻象”。为解决此问题且不采用计算成本高昂的替代方案,我们开展了一项大规模实证研究,涵盖六个零成本查询性能预测(QPP)指标,跨越三个认知层级:语义匹配(BEIR)、逻辑推理(BRIGHT)和时间推理(TEMPO)。我们的核心发现是,关键改进在于放弃幅度,转而采用分数分布信号:这一转变带来的增益比分布性替代方案之间的差异高出5-10倍。特别是,分数差距($s_1 - s_k$)以及分数幅度与方差的实用改编(LSMV)在幅度型置信度几乎无判别力的场景下,将弃权(不执行)AUROC提升了最多0.16。这些方法无需额外推理、重新训练或延迟,使其成为已部署RAG系统中幅度阈值化的实用零成本替代方案。
英文摘要
Many production RAG systems implement retrieval abstention by thresholding raw similarity scores, implicitly treating score magnitude as a confidence signal. We demonstrate that this practice degrades systematically as queries require reasoning beyond semantic matching. Across 11 retrieval architectures and 28 datasets, neural retrievers consistently assign high similarity scores to semantically related but constraint-violating documents, causing magnitude-based thresholds to collapse toward near-random abstention performance on logical and temporal reasoning tasks---a failure we term the Magnitude Mirage. To address this without computationally expensive alternatives, we conduct a large-scale empirical study of six zero-cost Query Performance Prediction (QPP) metrics across three cognitive tiers: semantic matching (BEIR), logical reasoning (BRIGHT), and temporal reasoning (TEMPO). Our central finding is that the key improvement comes from abandoning magnitude in favor of score-distribution signals: the gain from this shift exceeds the differences among distributional alternatives by a factor of 5-10$\times$. In particular, Score Gap ($s_1 - s_k$) and a practical adaptation of Score Magnitude and Variance (LSMV) improve abstention AUROC by up to 0.16 in settings where magnitude-based confidence provides little discriminative power. These methods require no additional inference, retraining, or latency, making them a practical zero-cost replacement for magnitude thresholding in deployed RAG systems.
发表机构
- UNSW Sydney(新南威尔士大学)
- University of Innsbruck(因斯布鲁克大学)
机构由 AI 辅助整理,请以论文原文为准。