多跳检索中LLM重排序器的逐查询门控
Per-Query Gating of LLM Rerankers for Multi-Hop Retrieval
浏览论文内容
中文总结 AI 辅助
本研究提出一种基于预调用特征的逐查询门控机制,用于在多跳检索中跳过无益的LLM重排序调用,在保持覆盖率的同时节省约一半调用,并量化了有害跳过的代价与校准乐观偏差。
中文摘要 AI 辅助
LLM重排序器在基于图增强的稠密检索流水线(如HippoRAG2)之上,每1000次查询增加约0.2-0.3美元的成本,并增加约一秒的尾部延迟;在三个多跳基准上,它们在九个(数据集,K)组合中的七个上提高了最终跳top-K覆盖率,最高提升达34.8个百分点。我们研究是否可以通过学习到的逐查询门控,仅使用LLM调用前可用的特征(两个检索列表的27个得分和词汇统计特征,以及一个小型查询嵌入的PCA),并配备可执行的回退机制,来跳过那些无益的重排序调用。所有选择,包括回退和阈值,都在训练折内做出,并仅对保留查询应用一次;有害跳过(即重排序本可找到目标而回退未能找到的情况)与总体覆盖率一同报告。在2WikiMultiHopQA、MuSiQue和HotpotQA的九个组合上,门控以平均保留LastHop@K损失1.2个百分点的代价跳过了51%的调用;四个组合满足预先注册的1个百分点规则,有害跳过发生在八个组合中(190次有害对136次有益),而在高提升组合上,相同跳过率的随机门控损失2至11个百分点。第二条规则根据对Platt校准后的有害概率(校准后ECE为0.025,校准前为0.094)的预期有害跳过率预算,为每个组合设定阈值:在1个百分点预算下,门控以-0.8个百分点的损失跳过42%的调用,产生66次有害跳过,六个组合在1个百分点以内,但实际有害程度在六个组合中超过承诺(平均1.45对0.83个百分点),我们量化了这种选择乐观偏差;0.5个百分点预算实际产生约1个百分点的损失。有害概率已校准但判别力有限(AUC为0.16至0.70)。早期版本报告了73%的“无损”节省;该数字依赖于一个神谕回退和一个错误的MuSiQue目标,我们对此两者均有记录。
英文摘要
LLM rerankers add of the order of \$0.2-0.3 per 1,000 queries and about a second of tail latency on top of a graph-augmented dense pipeline such as HippoRAG2, and on three multi-hop benchmarks they improve final-hop top-K coverage on seven of nine (dataset, K) cells, by up to +34.8 pp. We ask whether a learned per-query gate can skip the reranker where it will not help, using only features available before the LLM call (27 score and lexical statistics of the two retrieval lists plus a PCA of a small query embedding) with an executable fallback. Every choice, including the fallback and the threshold, is made inside the training fold and applied once to held-out queries, and harmful skips (the rerank would have found the target, the fallback did not) are reported next to the aggregate coverage. Across nine cells on 2WikiMultiHopQA, MuSiQue and HotpotQA the gate skips 51% of calls at an average held-out LastHop@K cost of 1.2 pp; four cells meet a pre-registered 1 pp rule, harmful skips occur in eight (190 harmful against 136 beneficial), and a random gate at the same skip rate loses 2 to 11 pp on the high-lift cells. A second rule sets each cell's threshold from a pre-specified budget on the expected harmful-skip rate over Platt-calibrated harm probabilities (ECE 0.025 after calibration, 0.094 before): at a 1 pp budget the gate skips 42% at -0.8 pp with 66 harmful skips and six cells within 1 pp, but realised harm exceeds the promise in six cells (mean 1.45 vs 0.83 pp), a selection optimism we quantify; a 0.5 pp budget realises about 1 pp. The harm probabilities are calibrated but barely discriminative (AUC 0.16 to 0.70). An earlier version reported 73% "lossless" savings; that figure rested on an oracle fallback and a wrong MuSiQue target, and we document both.