发表机构
Fraunhofer Heinrich-Hertz-Institut; Technische Universität Berlin; The Berlin Institute for the Foundations of Learning and Data (BIFOLD)(弗劳恩霍夫海因里希·赫兹研究所; 柏林工业大学; 柏林学习与数据基础研究所(BIFOLD))
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对3D肺结节分割,发现基于熵的聚合偶然不确定性无法捕捉病理存在的病例级歧义,提出轻量级监督歧义头可更有效建模该歧义,为医疗图像分割的不确定性估计提供了改进方向。
AI 中文摘要
不确定性估计对于深度学习在医学图像分割中的安全临床部署至关重要,偶然不确定性理论上旨在捕捉不可约的数据歧义。然而,基于熵的度量是否反映临床有意义的歧义,即关于是否存在病理的病例级分歧,仍知之甚少。与大多数关注像素级边界分歧的现有工作不同,我们系统评估了偶然不确定性捕捉存在歧义的效果。我们的评估涵盖3D肺结节分割,涉及四种采用蒙特卡洛dropout和深度集成的架构,使用LIDC-IDRI和外部验证队列(LNDb)。我们发现,基于熵的不确定性图与边界噪声和微小标注变化一致,但对存在歧义的判别信号不足。相比之下,在冻结的分割特征上训练的轻量级监督歧义头,在所有架构、指标和两个队列中,显著优于所有基于熵聚合的基线,且与在分歧监督下明确建模歧义的方法(概率U-Net、标注者混淆3D-UNet)相当或更优。定性特征空间分析显示,存在歧义已编码在像素级训练网络的冻结编码器特征中,仅被分割输出及其熵聚合丢弃。我们的发现揭示了偶然不确定性的理论承诺与其实际行为之间的根本不匹配,并建议从业者在安全关键应用中不应依赖基于熵的不确定性作为临床歧义的替代指标。
英文摘要
Uncertainty estimation is critical for the safe clinical deployment of deep learning in medical image segmentation, with aleatoric uncertainty theoretically designed to capture irreducible data ambiguity. However, whether entropy-based measures reflect clinically meaningful ambiguity, i.e. case-level disagreement about whether a pathology is present at all, remains poorly understood. Contrary to most prior work, which focused on pixel-wise boundary disagreement, we systematically evaluate how well aleatoric uncertainty captures presence ambiguity. Our evaluation spans 3D lung nodule segmentation across four architectures with Monte Carlo dropout and deep ensembles, on LIDC-IDRI and an external validation cohort (LNDb). We find that entropy-based uncertainty maps align with boundary noise and minor drawing variation but carry insufficient discriminative signal for presence ambiguity. In contrast, a lightweight supervised ambiguity head trained on frozen segmentation features substantially outperforms all entropy-aggregation-based baselines across architectures, metrics, and both cohorts, and matches or exceeds methods that explicitly model ambiguity under disagreement supervision (Probabilistic U-Net, Annotator-Confusion 3D-UNet). A qualitative feature-space analysis shows that presence ambiguity is already encoded in the frozen encoder features of pixel-wise-trained networks, only to be discarded by the segmentation output and its entropy aggregation. Our findings expose a fundamental mismatch between the theoretical promise of aleatoric uncertainty and its practical behavior, and suggest that practitioners should not rely on entropy-based uncertainty as a proxy for clinical ambiguity in safety-critical applications.