发表机构
University of Colorado Denver; Creighton University(科罗拉多大学丹佛分校; 克里顿大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究在临床图像退化下评估医学分割模型的不确定性,提出可重现框架,用合成脑肿瘤MRI队列训练U-Net等模型,经多种损坏类型及不同严重程度测试,结果表明不确定性感知推理可作安全层,还开源相关代码、模型和协议。
AI 中文摘要
医学图像分割模型在理想成像条件下常报告高基准精度,但在临床退化情况下可能失效,如传感器噪声、患者运动等因素会改变模型行为却无明显警告。我们提出一个用于在可控临床退化下评估不确定性感知分割的可重现框架。实验使用遵循BraTS协议的生物物理体模模拟器生成的合成多模态脑肿瘤MRI队列。我们训练U-Net和注意力U-Net基线进行多类肿瘤子区域分割,并通过蒙特卡洛随机失活增强模型以估计体素不确定性。在五种严重程度的八种临床相关损坏类型下,我们测量分割精度、校准、故障检测和选择性预测覆盖率。结果表明预测不确定性随退化增加并跟踪分割误差,可用于标记故障,不确定性感知推理可作为医生参与的放射学工作流程中的实用安全层。我们还发布了代码、训练模型和评估协议以支持直接重现。
英文摘要
Medical image segmentation models often report high benchmark accuracy under ideal imaging conditions, yet their failures under clinical degradation can be quiet: sensor noise, patient motion, low- resolution acquisition, and contrast variability may all alter model behavior without producing an obvious warning. We present a reproducible framework for evaluating uncertainty-aware segmentation under con- trolled clinical degradation. Our experiments use a synthetic multimodal brain tumor MRI cohort generated with a biophysical phantom simulator that follows the BraTS protocol. We train U-Net and Attention U-Net baselines for multi-class tumor sub-region segmentation and augment both models with Monte Carlo dropout to estimate per-voxel uncertainty. Across eight clinically motivated corruption types at five severity levels, we measure segmentation accuracy, calibration, failure detection, and selective prediction coverage. On clean data, Attention U-Net achieves a whole-tumor Dice of 0.990; under severe Gaussian noise, its performance falls to 0.089. Predictive uncertainty rises with degradation and tracks segmentation error (Pearson r = 0.53 under severity-3 Gaussian noise), allowing us to flag failures with an AUROC of 0.843. These results argue for uncertainty-aware inference as a practical safety layer in physician-in-the-loop radiology workflows. We release the code, trained models, and evaluation protocol to support direct reproduction.