发表机构
Bernoulli Institute, University of Groningen(格罗宁根大学伯努利研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究针对异构皮肤病变数据集,对比五种UQ方法发现不确定性可跨方法识别困难样本,Deep Ensembles在校准等指标上最优,基于不确定性的选择性转诊可减少大量错误。
AI 中文摘要
皮肤病变分类器可能在最重要的病例上出现自信的错误,因此判断何时不应信任预测结果在临床上与预测本身同样有用。我们研究了从多个ISIC来源汇集的数据集上的不确定性量化,该数据集具有共享主干和两个联合学习的头:一个二元良恶性分类头和一个五分类诊断头。我们在准确率、校准、不确定性分解和风险-覆盖率方面比较了五种UQ方法(MC Dropout、DropConnect、Flipout、Deep Ensembles、DUQ)。困难样本在很大程度上与方法无关:即使是熵分布较窄的方法,对困难样本的排序也相同(样本间熵相关性为0.54至0.91)。方法的选择对校准和不确定性分解的影响更大,其中Deep Ensembles是明显的赢家,而对困难样本的识别影响较小。该排序效果足够好,将最不确定的病例弃权(不执行)可消除不成比例的错误,支持基于不确定性的选择性转诊,本研究仅在分布内对该策略进行了评估。
英文摘要
Skin lesion classifiers can be confidently wrong on the cases that matter most, so knowing when a prediction should not be trusted is clinically as useful as the prediction. We study uncertainty quantification on a dataset pooled from many ISIC sources, with a shared backbone and two jointly learned heads: a binary malignant versus non-malignant head and a five-class diagnostic head. Five UQ methods (MC Dropout, DropConnect, Flipout, Deep Ensembles, DUQ) are compared on accuracy, calibration, uncertainty decomposition, and risk-coverage. Difficulty is largely method-agnostic: even methods with narrow entropy distributions rank the same samples as hard (per-sample entropy correlations of $0.54$ to $0.91$). The choice of method matters more for calibration and uncertainty decomposition, where Deep Ensembles is the clear winner, than for finding difficult cases. The ranking is also good enough that deferring the most uncertain cases removes a disproportionate share of errors, supporting uncertainty-based selective referral, evaluated here in-distribution only.
Comments12 pages, 9 figures, UNSURE 2026 @ MICCAI camera ready