检索主导的抽取式问答中的内在序列似然置信度:两个预设负例及其归因边界
Intrinsic Sequence-Likelihood Confidence in Retrieval-Dominated Extractive QA: Two Pre-Specified Negatives, and What They Do and Do Not Attribute
- Korea Institute of Science and Technology Information (KISTI)(韩国科学技术信息研究院)
- University of Science and Technology (UST)(科学技术联合大学院大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本研究证明在检索主导的抽取式问答中,基于序列似然的置信度信号(AUC 0.65-0.81)无法带来实质增益,路由与蒸馏机制均失败,最终产出为两个预设负例及其依赖关系。
AI中文摘要:
在抽取式文档问答中,当问题由包含答案的段落生成时(使得检索能恢复任何模式组合所能达到的92-99.8%的性能,无论其绝对准确度如何),基于置信度的机制几乎没有增益空间。在专业领域语料上微调开放语言模型,得到的模型自身置信度是一个诱人的控制信号:它可决定哪些查询需要进一步适配,以及哪些答案值得信任。我们在运行前固定的标准下评估了这两种用途,涉及四个7-9B参数规模的模型家族,其适配使闭卷F1分数最多提升+0.03,而两种用途均失败:在四个家族上的蒸馏触发机制(在其预设的三步迁移预算下)以及单模型试点中的路由与弃权(不执行)策略均未成功。仅靠检索在我们测试的每个正确性标准下恢复了最佳组合准确率的92-99.8%,使路由机制无实质增益。序列似然信号相对于该模式而言不足——在注册标准下受试者工作特征曲线下面积为0.65-0.81——适配前后均如此,标量重校准未改变,令牌级温度重缩放也未能一致改善。更精细的诊断取决于正确性标准和答案长度;在三个可测试的适配组合中,选择器消融显示置信度项在任何种子下均无统计上显著的下游收益;在Gemma上,移除置信度项使选择器从失败转为通过两项注册标准。可用的产出是一组预设负例及其依赖关系的明确化。
英文摘要:
In extractive document question answering whose questions were generated from the passages that contain their answers -- so that retrieval recovers 92-99.8% of what any mode combination could reach, whatever its absolute accuracy -- confidence-driven mechanisms have little to gain. Fine-tuning an open language model on a specialized domain corpus yields a model whose own confidence is a tempting control signal: it could decide which queries warrant further adaptation, and which answers to trust. We evaluate both uses under criteria fixed before the runs were executed, across four 7-9B model families whose adaptation moved closed-book F1 by at most +0.03, and both fail: a distillation trigger on all four families, under its pre-specified three-step transfer budget, and a routing-and-abstention policy in its single-model pilot. Retrieval alone recovers 92-99.8% of best-case combined accuracy under every correctness criterion we test, leaving routers no meaningful gain. The sequence-likelihood signal is insufficient relative to that mode -- area under the receiver operating characteristic curve 0.65-0.81 under the registered criterion -- before adaptation as well as after, unchanged by scalar recalibration and not consistently improved by token-level temperature rescaling. And the finer diagnostics depend on the correctness criterion and on answer length; on the three adapted combinations where we could test it, selector ablations show no statistically detectable downstream benefit from the confidence term on any seed; on Gemma, removing it changes the selector from failing to passing both registered criteria. The usable product is a set of pre-specified negatives with their dependencies made explicit.