AI 中文总结
针对多标签分类中校准度量未充分探索且现有分箱方案有缺陷的问题,提出对正负标签分配赋予同等权重的新分箱方案,以提供可信的标签级校准误差估计,并揭示大模型多标签校准的开放挑战。
AI 中文摘要
决定是否信任自动预测的一个关键因素是其置信度分数,该分数应经过校准以匹配预测正确的实际概率。大多数置信度校准指标针对二分类或多分类任务,而多标签校准在很大程度上仍未得到充分探索。多标签分类任务,例如为临床笔记分配医学代码或确定新闻主题,通常由大量负样本(即不适用的标签)主导。我们表明,现有的用于计算标签级期望校准误差的分箱方案要么低估了误差,要么仅仅反映了标签频率,要么存在许多实例极少的箱。为了获得可信的标签级校准误差,我们提出了一种新的分箱方案,该方案对正标签和负标签分配赋予同等权重。我们的实证研究表明,与现有的分箱方案相比,我们的新方案在层次化多标签分类和极端多标签分类中能够产生有意义的校准误差估计。我们还表明,校准大型语言模型用于多标签预测的置信度分数是一个开放的挑战。我们详细的分析通过提供一种用于测量多标签分类中校准的可靠评估指标,为进一步研究奠定了基础。
英文摘要
A key factor in deciding whether to trust an automatic prediction is its confidence score, which should be calibrated to match the actual probability of the prediction being correct. Most confidence calibration metrics target binary or multi-class tasks, while multi-label calibration remains largely underexplored. Multi-label classification tasks, such as assigning medical codes to clinical notes or determining news topics, are usually dominated by a large number of negatives, i.e., labels that do not apply. We show that existing binning schemes to compute label-wise expected calibration error either underestimate the error, simply reflect label frequency, or suffer from many bins with very few instances. To achieve trustworthy label-wise calibration errors, we propose a new binning scheme that gives equal weight to positive and negative label assignments. Our empirical study demonstrates that in contrast to existing binning schemes, our new scheme results in meaningful estimates of calibration error in hierarchical and in extreme multi-label classification. We also show that calibrating confidence scores of large language models for multi-label predictions is an open challenge. Our detailed analysis lays the foundation for further research by providing a solid evaluation metric for measuring calibration in multi-label classification.