arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.10406cs.IRcs.LG

基于逐标签单调投影的相关性决策后校准可靠性重排序

Post-Calibration Reliability Reranking of Relevance Decisions via Label-wise Monotone Projection

Inwoo Tae, Yongjae Lee

首次发表
浏览论文内容

中文总结 AI 辅助

针对检索系统校准后仍存在的标签依赖可靠性差异,提出逐标签单调可靠性投影(MRP)方法,通过残余风险重排序提升重排序效果与弃权效用,同时保持准确率与ECE。

中文摘要 AI 辅助

网页搜索、商品搜索和问答检索系统通常会为每个查询-候选对分配相关性标签和置信度分数:相关性标签描述页面、商品或段落与查询的匹配程度,而置信度常指导下游使用或弃权(不执行)决策。事后校准因此成为必要,因为校准偏差会导致系统过度信任错误预测或不必要地推迟正确预测。但校准主要是将置信度与平均正确率对齐,无法消除同一校准置信水平内存在的、依赖于预测标签的可靠性差异。针对这一缺口,本文提出逐标签单调可靠性投影(MRP)方法,该方法学习逐标签单调函数,将校准后的置信度映射到正确性可靠性,同时保留原始预测标签和类别概率;生成的可靠性分数会根据残余风险对固定预测结果进行重排序。在六个信息访问相关性数据集和多个事后校准器上的实验显示,MRP可提升可靠性重排序效果与平均弃权(不执行)效用,同时保持全覆盖准确率与预期校准误差(ECE)。结构消融实验表明,主要增益来自逐标签残余可靠性而非全局置信度重映射。本文还分析了MRP可靠性分数可嵌入回最高标签概率几何的情况,指出该投影可用于兼容性分析,但与核心的可靠性重排序目标不同。相关代码将公开提供。

英文摘要

Web search, product search, and question-answering retrieval systems often assign a relevance label and confidence score to each query-candidate pair. The relevance label describes how well a page, product, or passage matches the query, while the confidence often guides downstream use or fallback decisions. Post-hoc calibration is therefore needed because misaligned confidence can make systems over-trust wrong predictions or unnecessarily defer correct ones. However, calibration mainly aligns confidence with average correctness, and does not remove predicted-label-dependent reliability differences that remain within the same calibrated confidence level. We address this gap with Label-wise Monotone Reliability Projection (MRP), which learns label-wise monotone functions that map calibrated confidence to correctness reliability while preserving the original predicted labels and class probabilities. The resulting reliability score reranks fixed predictions according to residual risk. Across six information access relevance datasets and multiple post-hoc calibrators, MRP improves reliability reranking and average fallback utility while preserving full-coverage accuracy and ECE. Structural ablations show that the main gains come from label-wise residual reliability rather than from global confidence remapping. We further analyze when MRP reliability scores can be embedded back into top-label probability geometry, showing that this projection is useful as a compatibility analysis but is distinct from the main reliability-reranking objective. The implementation will be made publicly available.

发表机构

  • UNIST(蔚山科学技术院)
  • LinqAlpha(林克阿尔法)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑