发表机构
University College Dublin(都柏林大学学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究通过多标注者视觉基准,发现模型不确定性及预测多重性与人类歧义对齐较弱,表明模型不确定性不宜作为高风险决策的默认可信信号。
AI 中文摘要
人类与模型的对齐对于可信的AI辅助决策系统至关重要。然而,大多数工作仅针对单一真实标签评估模型预测,忽视了人类自身在标签上常常存在分歧,这恰恰是真实歧义的一种信号。我们研究了模型是否在人类认为困难的相同实例上表现挣扎。我们在两个视觉数据集(FER+和CIFAR-10H)上对此进行测量,这两个数据集中每个图像都有多个人类标注,以捕捉人类的分歧模式。我们评估了三种架构(ResNet、EfficientNet、MobileNetV3)下的八个预训练模型,分两部分进行:首先,模型不确定性(softmax置信度、熵)是否与人类分歧相关;其次,预测多重性度量(模型间分歧、Jensen-Shannon散度)是否与之相关。我们发现事实并非如此:两个维度的对齐都很弱。在离散标签层面,CIFAR-10H中50.4%的图像和FER+中33.5%的图像从人类那里获得了多个有效分类,而模型仅收敛于一个。这些实例代表了一个关键失败案例:人类感知到歧义并会请求专家审查,而模型却自信地做出决定。在连续分数层面,单一模型不确定性与人类分歧的相关性较弱(ρ = 0.24–0.55),预测多重性仅带来适度改进。广泛使用的不确定性量化方法无法可靠地识别人类认为歧义的实例。在高风险场景的决策中,模型不确定性不应默认被视为可信信号。
英文摘要
Human-model alignment is critical for trustworthy AI-assisted decision-making systems. Yet, most work evaluates model predictions against single ground-truth labels, overlooking that humans themselves often disagree on labels, a signal of genuine ambiguity. We investigate whether models struggle on the same instances that humans find difficult. We measure this on two vision datasets (FER+ and CIFAR-10H) where multiple human annotations per image capture human disagreement patterns. We evaluate eight pretrained models across three architectures (ResNet, EfficientNet, MobileNetV3) in two parts: first, whether model uncertainty (softmax confidence, entropy) correlates with human disagreement, and second, whether predictive multiplicity measures (inter-model disagreement, Jensen-Shannon divergence) do. We find that it does not: alignment is weak in both dimensions. At the discrete label level, 50.4% of CIFAR-10H images and 33.5% of FER+ images receive multiple valid classifications from humans, while the models converge on only one. These instances represent a critical failure case where humans perceive ambiguity and would request expert review, yet models decide confidently. At the continuous score level, single-model uncertainty correlates weakly with human disagreement ($ρ= 0.24--0.55$), and predictive multiplicity provides only modest improvement. Widely-used uncertainty quantification methods do not reliably identify instances humans find ambiguous. Model uncertainty should not be treated as a trustworthy signal by default for decision-making in high-stakes scenarios.