用于与架构无关的脑肿瘤MRI分类的蒙特卡罗随机失活不确定性和熵阈值选择性预测
Monte Carlo Dropout Uncertainty and Entropy-Thresholded Selective Prediction for Architecture-Agnostic Brain Tumor MRI Triage
AI总结:
该研究针对脑肿瘤MRI分类,提出基于蒙特卡罗随机失活读取不确定性并转化为熵阈值规则的流程,通过特定图像分区和多模型多指标评估,结果表明此不确定性流程有效,可校准且能提高分类准确性,提供架构无关的脑肿瘤MRI分类基础。
AI中文摘要:
深度网络现在在MRI上对脑肿瘤进行亚型分类的能力与专业读者相当,但准确性并非其无法应用于临床的原因。在护理点重要的是模型的置信度是否可信赖,以标记其可能误分类的病例并将其提交给人类判断。确定性估计无法做到这一点:与分类器一起训练的辅助置信度头会崩溃为几乎恒定的输出,无法说明正确性。本研究为四类脑肿瘤MRI(胶质瘤、脑膜瘤、垂体瘤、无肿瘤)提出了一种以不确定性为先的流程,该流程通过在T = 20次传递上的蒙特卡罗(MC)随机失活读取预测不确定性,并将所得熵转化为将不确定病例提交给放射科医生的明确规则。我们通过感知哈希聚类对7200张图像进行分区,消除了在简单分割下会夸大准确性的近重复泄漏,并在ViT - B/16和ResNet - 50上沿三个轴(辨别力、校准和选择性预测)对五个种子进行了评估。两者都具有很强的辨别力(宏AUC为0.994;准确率为0.962和0.964),且没有种子能将它们区分开(5个中0个显著,p < 0.05),所以结果由不确定性流程而非网络驱动。单个温度标量将确定性softmax调整到紧密校准(预期校准误差为0.016 - 0.020),推迟最不确定的5%的病例可将其余病例的准确率提高到约0.98(风险覆盖曲线下面积为0.010 - 0.011)。因此,MC随机失活不确定性在这里经过校准、不会崩溃,并且可以通过具体的推迟规则直接采取行动,为内部验证下的校准、提交给人类的脑肿瘤MRI分类提供了与架构无关的基础。
英文摘要:
Deep networks now subtype brain tumors on MRI about as well as specialist readers, yet accuracy is not what keeps them out of the clinic. What matters at the point of care is whether a model's confidence can be trusted to flag the cases it is likely to misclassify and defer them to a human. Deterministic estimates cannot: an auxiliary confidence head trained alongside the classifier collapses to a near-constant output that says nothing about correctness. This study proposes an uncertainty-first pipeline for four-class brain tumor MRI (glioma, meningioma, pituitary, no tumor) that reads predictive uncertainty from Monte Carlo (MC) Dropout over T = 20 passes and turns the resulting entropy into an explicit rule for deferring uncertain cases to a radiologist. We partitioned 7,200 images by perceptual-hash cluster, closing the near-duplicate leakage that inflates accuracy under naive splitting, and evaluated the pipeline on ViT-B/16 and ResNet-50 across five seeds along three axes: discrimination, calibration, and selective prediction. Both discriminate strongly (macro-AUC 0.994; accuracy 0.962 and 0.964), and no seed separates them (0 of 5 significant, p < 0.05), so the result is driven by the uncertainty pipeline, not the network. A single temperature scalar pulls the deterministic softmax into tight calibration (expected calibration error 0.016-0.020), and deferring the most uncertain 5% of cases lifts accuracy on the rest to about 0.98 on both (area under the risk-coverage curve 0.010-0.011). MC-Dropout uncertainty here is thus calibrated, non-collapsing, and directly actionable through a concrete deferral rule, providing an architecture-agnostic basis for calibrated, defer-to-human brain tumor MRI triage under internal validation.