arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

温度缩放还不够:人类标签分布下的校准差距

Temperature Scaling Is Not Enough: Calibration Gaps Under Human Label Distributions

Wisdom Dogah

arXiv 2607.13423首次发表:更新:

发表机构

Faculty of Computing and Mathematical Sciences, University of Mines and Technology (UMaT); BlackMatrix AI Research(矿业与技术大学(UMaT)计算与数学科学学院; 黑矩阵人工智能研究)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究在违反真实标签假设时温度缩放的校准特性及与模型规模的关系,用CIFAR-10H和ChaosNLI数据集评估,发现温度缩放校准硬标签不如直接校准软标签,差距随模型规模在部分领域增长,多类保序回归也有类似结论,揭示相关校准协议问题。

AI 中文摘要

温度缩放是现代深度学习中主要的事后校准方法。其理论依据基于一个很少明确说明的假设:真实标签是独热且确定的。实际上,标签常常是软的、众包的或真正分布性的,反映了人类标注者之间的真实分歧而非标注噪声。我们研究当这一假设被违反时温度缩放是否仍保持其校准特性,以及由此产生的性能下降是否取决于模型规模。使用CIFAR-10H和ChaosNLI这两个具有人类标注软标签分布的公开数据集,我们在硬独热和软分布标签目标下评估了每个模态的三种模型规模。在所有九种配置中,我们发现了一个正向软标签校准差距:在硬标签上校准的温度缩放始终不如直接在软标签上校准的最优模型,布里尔得分差距在0.002到0.134之间。在视觉领域以及ChaosNLI的SNLI衍生分割上,差距随模型规模单调增长,并且在语言领域(平均差距0.079)比在视觉领域(平均差距0.003)大得多。在匹配域训练后,MNLI衍生分割上的规模排序反转仍然存在;我们认为这对于规模假设尚无定论,主要归因于该分割上接近随机的准确率。作为第二个事后基线,多类保序回归得出了相同的定性结论:在所有九种配置中都存在正向软标签差距,并且语言领域的差距比视觉领域更大。这些发现表明,基于多数投票标签构建的校准协议在标签模糊是结构性的任何地方都会系统性地错误陈述模型可靠性,这对安全关键环境中的部署有直接影响。

英文摘要

Temperature scaling is the dominant post-hoc calibration method in modern deep learning. Its theoretical justification rests on an assumption that is rarely stated explicitly: that ground-truth labels are one-hot and deterministic. In practice, labels are frequently soft, crowd-sourced, or genuinely distributional, reflecting real disagreement among human annotators rather than annotation noise. We study whether temperature scaling retains its calibration properties when this assumption is violated, and whether any resulting degradation depends on model scale. Using CIFAR-10H and ChaosNLI, two publicly available datasets with human-annotated soft label distributions, we evaluate three model scales per modality under both hard one-hot and soft distributional label targets. Across all nine configurations we find a positive soft-label calibration gap: temperature scaling calibrated on hard labels consistently underperforms an oracle calibrated directly on soft labels, with Brier Score gaps ranging from 0.002 to 0.134. The gap grows monotonically with model scale in the vision domain and on the SNLI-derived split of ChaosNLI, and is substantially larger in the language domain (mean gap 0.079) than in vision (mean gap 0.003). A scale-ordering reversal on the MNLI-derived split remains after matched-domain training; we treat it as inconclusive for the scale hypothesis and attribute it primarily to near-chance accuracy on that split. As a second post-hoc baseline, multiclass isotonic regression yields the same qualitative conclusion: positive soft-label gaps in all nine configurations, and larger gaps in language than in vision. These findings suggest that calibration protocols built on majority-vote labels systematically misstate model reliability wherever label ambiguity is structural, with direct consequences for deployment in safety-critical settings.

Comments9 pages. Code and per-seed results: https://github.com/dogahwisdom/temperature-scaling-research

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑