多跳检索中的可预测失败:分数分布置信度评分与弃权(不执行)
Predictable Failure in Multi-Hop Retrieval: Score-Distributional Confidence Scoring and Abstention
AI总结:
本研究证明多跳检索失败在结构上可预测,提出RegimeAbstain方法,利用检索置信度分数(RCS)实现校准弃权策略,在多个基准上显著降低自信错误答案率(CWAR),并验证了特征机制的领域无关性。
AI中文摘要:
多跳检索失败并非在查询间均匀分布:它们聚集在结构上可预测的子群体中。我们证明了两个结果,将这些结构形式化。第一(CWAR可约性):当且仅当检索特征携带关于成功的信息时,自信失败减少是可实现的,这一条件由LLM评判流水线满足,但在仅稠密设置中显著较弱,解释了两种机制之间AUC-AC的差距。第二(特征机制互补性):没有单一的ANN分数特征在所有失败机制中达到最佳预测性能;主导特征在不同数据集间不同(MuSiQue上为查询长度,HoVer上为hop-1集中度),且一个构造性见证对显示每个特征在一种机制中是必要的,在另一种机制中无贡献。我们将这些原则实例化于RegimeAbstain中,它计算一个检索置信度分数(RCS),这是多达九个查询-ANN结构特征的逻辑函数,所有特征无需额外LLM调用即可获得,并使用它实施校准的弃权(不执行)策略。我们定义了自信错误答案率(CWAR)指标,并在三个多跳基准(MuSiQue、2WikiMultiHopQA、HoVer)和两种检索架构(LLM评判和仅稠密)上评估,覆盖五种失败机制,CWAR从14.5%到62.1%。RCS在所有五种条件下对八个置信度基线达到最佳或共同最佳AUC-AC。在MuSiQue(LLM评判)上,RCS在50%覆盖率下将CWAR从39.5%降至20.6%(相对减少47.8%),ECE=0.035。在MuSiQue上训练的模型迁移到2WikiMultiHopQA,仅损失-0.5个百分点AUC,确认了机制特征的领域无关结构。
英文摘要:
Multi-hop retrieval failures are not uniformly distributed across queries: they cluster in structurally predictable subpopulations. We prove two results formalizing this structure. First (CWAR Reducibility): confident-failure reduction is achievable if and only if retrieval features carry mutual information about success, a condition satisfied by LLM-judge pipelines but substantially weaker in dense-only settings, explaining the AUC-AC gap between regimes. Second (Feature Regime Complementarity): no single ANN score feature achieves best predictive performance across all failure regimes; the dominant feature differs between datasets (query length on MuSiQue, hop-1 concentration on HoVer), and a constructive witness pair shows each is necessary in one regime and non-contributory in the other. We instantiate these principles in RegimeAbstain, which computes a Retrieval Confidence Score (RCS), a logistic function of up to nine query-ANN structural features, all available without any additional LLM call, and uses it to implement a calibrated abstention policy. We define the Confident-Wrong-Answer Rate (CWAR) metric and evaluate across three multi-hop benchmarks (MuSiQue, 2WikiMultiHopQA, HoVer) and two retrieval architectures (LLM-judge and dense-only), covering five failure regimes with CWAR from 14.5% to 62.1%. RCS achieves best or co-best AUC-AC in all five conditions against eight confidence baselines. On MuSiQue (LLM-judge), RCS reduces CWAR from 39.5% to 20.6% at 50% coverage (47.8% relative reduction), with ECE=0.035. A model trained on MuSiQue transfers to 2WikiMultiHopQA with only -0.5pp AUC loss, confirming the domain-agnostic structure of regime features.