发表机构
Beihang University(北京航空航天大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究在知识密集型问答中,利用黑盒语言模型的置信度指导检索路由。核心方法是BeyondUncertainty,先得临时答案和置信度,依阈值决定检索策略。该方法提升了平均词元级F1,减少检索段落,在多数设置中优于随机分配,揭示了证据获取与词元效率的权衡。
AI 中文摘要
检索增强生成可改善知识密集型问答,但随意检索会引入无关证据和不必要计算。我们研究黑盒语言模型中表达出的置信度能否作为检索路由的可行信号。我们的方法BeyondUncertainty先得出结构化临时答案和置信度估计,再应用在验证数据上选择并在测试评估前冻结的特定模型阈值。低置信度问题进行前5的TF-IDF检索并再次调用答案,高置信度问题直接返回临时答案。我们在六个问答基准、三个模型家族和三种检索策略上评估了27000个策略实例。BeyondUncertainty实现了0.483的平均词元级F1,相比总是检索的0.467和不检索的0.401有所提升,同时检索段落减少20.4%。在每个数据集-模型单元内按检索问题数量匹配时,它在18个设置中的17个优于事后随机分配,平均F1增益为0.024。虽然作为绝对概率校准不佳,但探测不确定性适度预测了问题级检索收益(AUROC = 0.628)。然而,额外的探测使总词元使用增加28.2%,揭示了更有选择性的证据获取与端到端词元效率之间的权衡。
英文摘要
Retrieval-augmented generation improves knowledge-intensive question answering, but indiscriminate retrieval can introduce irrelevant evidence and unnecessary computation. We investigate whether verbalized confidence from black-box language models can serve as an actionable signal for retrieval routing. Our method, BeyondUncertainty, first elicits a structured provisional answer and confidence estimate, then applies a model-specific threshold selected on held-out validation data and frozen before test evaluation. Low-confidence questions receive top-5 TF-IDF retrieval followed by a second answer call, whereas high-confidence questions return the provisional answer directly. We evaluate 27,000 policy instances across six QA benchmarks, three model families, and three retrieval policies. BeyondUncertainty achieves 0.483 mean token-level F1, compared with 0.467 for always retrieval and 0.401 for no retrieval, while reducing retrieved passages by 20.4\% relative to always retrieval. When matched on the number of questions routed to retrieval within each dataset-model cell, it outperforms a post-hoc random allocation in 17 of 18 settings, with an average gain of 0.024 F1. Although poorly calibrated as an absolute probability, probe uncertainty modestly predicts question-level retrieval benefit (AUROC = 0.628). However, the additional probe increases total token usage by 28.2\%, revealing a trade-off between more selective evidence acquisition and end-to-end token efficiency.
Comments9 pages, 6 figures, 6 tables