巴氏涂片分类的选择性预测与不确定性感知转诊
Selective Prediction and Uncertainty-Aware Referral for Pap Smear Classification
浏览论文内容
中文总结 AI 辅助
本研究在Herlev巴氏涂片数据集上比较了单个轻量级Transformer与软投票集成在选择性预测中的表现,发现集成显著降低AURC并扩展零错误覆盖率,但校准更差,区分了排序可信度与置信度准确性两个性质。
中文摘要 AI 辅助
用于宫颈细胞学的深度学习模型几乎总是在假设每个预测都必须被执行的情况下进行评估,然而,与细胞病理学家一同部署的筛查系统并不需要对每张涂片进行分类:它可以推迟(转诊)其最不确定的病例。评估这样的系统不仅需要询问其正确的频率,还需要询问其置信度是否将其错误排在底部。本文在Herlev巴氏涂片数据集上,以二分类正常与异常的形式,研究了选择性预测和不确定性感知转诊。两个轻量级Transformer骨干网络(Swin-Tiny、TinyViT-5M)从ImageNet预训练权重出发,使用加权随机采样在Herlev上进行了微调,并通过在留出的校准子集上拟合的事后温度缩放进行校准,并与两个模型的软投票集成进行了比较。判别性能与预期校准误差(ECE)一起报告,主要终点是风险覆盖率曲线下面积(AURC)。两种配置在准确率或宏F1上未检测到统计学显著差异,然而集成将AURC减半(0.0022对比0.0045,降低51.8%,在所有五个折中均更低),并将零错误覆盖率从合并测试预测的18.3%扩展到72.8%。然而,同一集成在绝对项上校准更差(ECE 0.0339对比0.0247),并产生更多假阴性(14对比10)。这些结果区分了两个经常被混淆的性质:按可信度对预测排序的能力,以及置信度值本身的准确性。集成改善了前者而降低了后者,前者直接控制观察到的风险覆盖率权衡,而后者控制所报告置信度值的解释。
英文摘要
Deep learning models for cervical cytology are almost always evaluated as if every prediction must be acted upon, yet a screening system deployed alongside a cytopathologist need not classify every slide: it can defer the cases it is least certain about. Evaluating such a system requires asking not only how often it is correct, but whether its confidence ranks its errors to the bottom. This paper studies selective prediction and uncertainty-aware referral on the Herlev Pap smear dataset under a binary Normal-versus-Abnormal formulation. Two lightweight transformer backbones (Swin-Tiny, TinyViT-5M) are fine-tuned on Herlev from ImageNet-pretrained weights with weighted random sampling, calibrated by post-hoc temperature scaling fit on a held-out calibration subset, and compared against a soft-voting ensemble of both models. Discrimination is reported alongside expected calibration error (ECE) and, as the primary endpoint, the area under the risk-coverage curve (AURC). No statistically significant difference was detected between the two configurations in accuracy or macro-F1, yet the ensemble halves AURC (0.0022 vs. 0.0045, a 51.8% reduction, lower in all five folds) and extends the coverage at which zero errors are made from 18.3% to 72.8% of the pooled test predictions. The same ensemble is nonetheless worse calibrated in absolute terms (ECE 0.0339 vs. 0.0247) and produces more false negatives (14 vs. 10). These results separate two properties that are frequently conflated: the ability to rank predictions by trustworthiness, and the accuracy of the confidence values themselves. Ensembling improves the former while degrading the latter, and the former directly governs the observed risk-coverage tradeoff, whereas the latter governs the interpretation of the reported confidence values.
发表机构
- Binghamton University(宾汉姆顿大学)
机构由 AI 辅助整理,请以论文原文为准。