发表机构
Erasmus MC, University Medical Center Rotterdam; Amsterdam UMC, University of Amsterdam; GE HealthCare; Erasmus MC Cancer Institute; Medical Delta(鹿特丹大学医学中心伊拉斯谟MC; 阿姆斯特丹大学医学中心阿姆斯特丹UMC; GE医疗; 伊拉斯谟MC癌症研究所; 医疗三角洲)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究在多任务胶质瘤诊断框架中评估不确定性量化方法,比较MCD、深度集成及其组合,发现中等丢弃率校准最佳,各方法无一致优势,为可信赖AI提供任务感知评估策略。
AI 中文摘要
不确定性量化(UQ)是高风险医学图像分析中可信赖人工智能的关键要求。在本工作中,我们在一个基于MRI的胶质瘤诊断多任务深度学习框架中评估了UQ,该框架执行肿瘤分割并预测IDH突变状态、1p/19q共缺失状态和肿瘤分级。采用蒙特卡洛丢弃法(MCD)进行详细的面向任务的预测性、偶然性和认知不确定性分析。我们评估了MC样本收敛性、校准、错误检测、选择性预测、与分割性能的关联,以及体素级不确定性聚合对病例级可靠性的影响。我们还将MCD与深度集成(DE)和蒙特卡洛深度集成(MCDE)进行比较,考察分割质量与分类之间的交互作用,并评估一个整合分割和分类不确定性的复合信任分数。在各任务中,不确定性估计支持有意义的错误检测,而校准依赖于丢弃率,中等丢弃率产生最可靠的概率。不确定性分解提供了任务相关的可解释性,但并未始终优于单独使用预测性不确定性进行错误检测。DE和MCDE显示出相当的实用效用,没有一种方法在所有任务和指标上持续占优。复合信任分数在选择性预测方面并未持续优于分类不确定性。总体而言,我们的结果为胶质瘤诊断可信赖人工智能的开发提供了面向任务的评估策略和实用指导。
英文摘要
Uncertainty Quantification (UQ) is a key requirement for trustworthy AI in high-stakes medical image analysis. In this work, we evaluate UQ in a multi-task Deep Learning framework for MRI-based glioma diagnosis that performs tumor segmentation and predicts IDH mutation status, 1p/19q co-deletion status, and tumor grade. Monte Carlo Dropout (MCD) is used for a detailed task-aware analysis of predictive, aleatoric, and epistemic uncertainty. We assess MC sample convergence, calibration, error detection, selective prediction, associations with segmentation performance, and the effect of voxel-wise uncertainty aggregation on case-level reliability. We also compare MCD with Deep Ensembles (DE) and Monte Carlo Deep Ensembles (MCDE), examine interactions between segmentation quality and classification, and evaluate a composite trust score integrating segmentation and classification uncertainty. Across tasks, uncertainty estimates supported meaningful error detection, while calibration depended on the dropout rate, with moderate rates yielding the most reliable probabilities. Uncertainty decomposition provided task-dependent interpretability but did not consistently improve error detection over predictive uncertainty alone. DE and MCDE showed comparable operational utility, with no method consistently dominating across tasks and metrics. The composite trust score did not consistently outperform classification uncertainty for selective prediction. Overall, our results provide a task-aware evaluation strategy and practical guidance for the development of trustworthy AI for glioma diagnosis.
CommentsAccepted for publication at the Journal of Machine Learning for Biomedical Imaging (MELBA) https://melba-journal.org/2026:033
Journal refMachine.Learning.for.Biomedical.Imaging. 2026 (2026)
DOI:10.59275/j.melba.2026-456d