发表机构
The University of Manchester(曼彻斯特大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出不确定性驱动训练框架,利用MCD和EDL估计指导损失重加权,在多种骨干网络上提升CT肺结节分类的校准性能,ECE降低高达65%。
AI 中文摘要
在本工作中,我们提出了一种用于三维计算机断层扫描(CT)肺结节分类的不确定性驱动训练框架,其中基于验证集的不确定性估计指导损失重新加权,以增强预测性能和概率校准。我们考虑了两种不确定性量化(UQ)方法:蒙特卡洛丢弃(MCD)和证据深度学习(EDL)。两者均提供逐类不确定性估计,用于调节损失并鼓励关注困难或不可靠的类别。该框架使用ResNet、DenseNet、EfficientNet、视觉变换器(ViT)和Swin变换器骨干网络在两个数据集上进行了评估:临床LIDC-IDRI队列和NoduleMNIST3D基准。不确定性驱动训练实现了与传统训练相似的分类性能,同时显著改善了校准,在LIDC-IDRI上预期校准误差(ECE)降低了高达65%。EDL在较浅架构上以单次推理获得具有竞争力的性能,而MCD在更深网络上更为稳健。跨架构家族的分析表明,不确定性驱动训练对卷积骨干网络的益处比基于变换器的架构更一致:特别是EDL在ViT上性能下降,这表明狄利克雷证据参数化可能在较低输入分辨率下与基于注意力的架构产生不利交互。后验温度缩放在所有配置中被证明非常有效,表明即使没有显式的不确定性感知训练,简单的标量校准也能具有竞争力。我们的结果表明,将UQ集成到训练循环中可以显著改善概率校准,并支持三维医学成像模型更可信的部署。
英文摘要
In this work, we present an uncertainty-driven training framework for three-dimensional computed tomography (CT) lung nodule classification, where validation-based uncertainty estimates guide loss reweighting to enhance predictive performance and probability calibration. Two Uncertainty Quantification (UQ) methods are considered: Monte Carlo Dropout (MCD) and Evidential Deep Learning (EDL). Both provide per-class uncertainty estimates that modulate the loss and encourage focus on hard or unreliable classes. The framework is evaluated with ResNet, DenseNet, EfficientNet, Vision Transformer (ViT), and Swin Transformer backbones on two datasets: the clinical LIDC-IDRI cohort and the NoduleMNIST3D benchmark. Uncertainty-driven training achieves classification performance similar to conventional training while substantially improving calibration, with an expected calibration error (ECE) reduced by up to 65% on LIDC-IDRI. EDL attains competitive performance on shallower architectures with single-pass inference, whereas MCD is more robust on deeper networks. Analysis across architectural families reveals that uncertainty-driven training benefits convolutional backbones more consistently than transformer-based architectures: EDL in particular degrades on ViT, suggesting that the Dirichlet evidence parameterisation may interact unfavourably with attention-based architectures at lower input resolutions. A posteriori temperature scaling proves highly effective across all configurations, indicating that a simple scalar calibration can be competitive even without explicit uncertainty-aware training. Our results indicate that integrating UQ into the training loop can significantly improve probabilistic calibration and support more trustworthy deployment of three-dimensional medical imaging models.
Journal refArtificial Neural Networks and Machine Learning ICANN 2026
DOI:10.1007/978-3-032-38407-2_21