发表机构
University of Birmingham(伯明翰大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究针对BraTS-GoAT分割任务,对比nnU-Net单模型与深度集成模型的可靠性,发现集成模型的分歧是更敏感的采集偏移指标,其贡献为严谨的不确定性方法可靠性比较。
AI 中文摘要
深度网络在分布内可准确分割脑肿瘤,但当输入与训练数据存在差异时可能发生静默失效,该风险是临床部署的核心问题,也是BraTS-GoAT泛化任务的前提。本研究不仅关注模型的分割性能,还探究其不确定性是否能识别自身的错误。在BraTS-GoAT(任务3)上,我们训练了5折交叉验证的nnU-Net基线(每个病例保留1个预测)和3个种子的深度集成模型,两者均基于与区域相关的掩码按病例聚合,评估校准度和错误检测能力。分布内场景下,3个种子的集成模型在已有的强单模型保留拆分基础上,校准度提升最为明显;分布偏移场景下,该分离效应显现。在使用分级合成损坏作为采集偏移代理的受控鲁棒性研究中,单模型的置信度保持平稳,但其准确率和校准度下降;而成员间的分歧则急剧上升,较干净条件高出约四分之一至三分之一,是单模型响应的数倍。在官方验证排行榜上,这些折的5折集成模型取得全肿瘤Dice为0.87的成绩。泛化差距集中在更难的区域,表现为对未见队列中小卫星病灶的特征性漏检。在合成研究中,3个种子成员间的分歧是比单模型置信度更敏感的病例级采集偏移指标,但随着损坏严重程度增加,其体素级错误定位能力减弱。本研究的贡献是开展了严谨、客观的可靠性比较,而非宣称某一不确定性方法具有主导性。
英文摘要
Deep networks segment brain tumours accurately in-distribution, but can fail silently when the input differs from their training data. That risk is central to clinical deployment and is the premise of the BraTS-GoAT generalizability task. We ask not only how well a model segments, but whether its uncertainty knows when it is wrong. On BraTS-GoAT (Task 3) we train a 5-fold cross-validated nnU-Net baseline (one held-out prediction per case) and a 3-seed deep ensemble. Both are evaluated for calibration and error detection on a per-region relevant mask, aggregated per case. In-distribution the 3-seed ensemble improves modestly over the already strong single model on the same held-out split, with the clearest gain in calibration. The separation appears under shift. In a controlled robustness study using graded synthetic corruptions as a proxy for acquisition shift, the single model's confidence stays flat while its accuracy and calibration degrade. Inter-member disagreement instead rises steeply, about a quarter to a third above the clean condition, several times the single model's response. On the official validation leaderboard the 5-fold ensemble of those folds attains whole-tumour Dice 0.87. The generalization gap is concentrated on the harder regions, with a characteristic failure of missing small, satellite lesions on unseen cohorts. In the synthetic study, disagreement among the 3-seed members is a more sensitive case-level indicator of acquisition shift than single-model confidence. Its per-voxel error localisation weakens as severity grows. The contribution is a rigorous, honest reliability comparison rather than a claim that any one uncertainty method dominates.
Comments12 pages, 4 figures, 2 tables. Conditionally accepted at MICCAI 2026 BraTS-GoAT challenge workshop. Code: https://github.com/riyashet-hds/brats-goat-reliability