基于扩散模型的无监督解剖特征学习:用去噪扩散概率模型增强医学图像分割
Unsupervised Anatomical Feature Learning via Diffusion Models: Enhanced Medical Image Segmentation with Denoising Diffusion Probabilistic Models
浏览论文内容
中文总结 AI 辅助
该研究提出用无监督DDPM预训练提取解剖特征,将编码器权重迁移至U-Net实现医学图像分割,在BTCV数据集上显著提升多器官分割性能,低数据场景下仍保持稳健表现。
中文摘要 AI 辅助
医学图像分割的像素级标注获取是一个严重瓶颈。传统U-Net架构虽有效,但仅学习局部纹理模式,缺乏全局解剖结构感知,在数据不足时会导致边界勾画失败。本研究提出利用无监督去噪扩散概率模型(DDPMs)提取解剖特征:在21例未标注腹部CT扫描上训练DDPM以学习结构表示,将编码器权重迁移至下游分割任务,在BTCV多器官数据集上评估。扩散预训练显著提升了肝脏分割性能:Dice系数从0.75±0.36升至0.93±0.16(p<5.33×10^-26,Cohen's d=0.529),平均表面距离(ASD)降低66%,95百分位豪斯多夫距离(HD95)减少45%;肾脏分割的Dice系数从0.90±0.19升至0.95±0.10(p<4.01×10^-11)。多器官合并性能显示方差降低68%,边界精度提升74%(Dice=0.95±0.07)。关键的是,冻结编码器模型在未接触分割标注的情况下保留了超过80%的微调性能,证明存在已学习的解剖先验。在低数据场景中,经扩散预训练的模型仅用50%标注数据(肝脏Dice=0.92,肾脏Dice=0.94)、25%甚至10%标注数据(肝脏Dice=0.89,肾脏Dice=0.71)仍保持稳健性能。利用未标注图像进行基于扩散的预训练,可在人工监督前成功嵌入稳健的解剖特征先验,将U-Net转化为具备解剖感知能力的系统。
英文摘要
Acquiring pixel-level annotations for medical image segmentation is a severe bottleneck. Traditional U-Net architectures, while effective, learn local texture patterns and lack awareness of global anatomical structures, leading to boundary delineation failures in low-data regimes. This research paper proposes utilizing unsupervised Denoising Diffusion Probabilistic Models (DDPMs) to extract anatomical features. We train a DDPM on 21 unlabeled abdominal CT scans to learn structural representations, transferring the encoder weights to a downstream segmentation task evaluated on the BTCV multi-organ dataset. Diffusion pretraining significantly improved liver segmentation: Dice increased from $0.75\pm0.36$ to $0.93\pm0.16$ ($p < 5.33\times10^{-26}$, 0.529 Cohen's d), Average Surface Distance (ASD) decreased by 66%, and 95th-percentile Hausdorff Distance (HD95) reduced by 45%. For kidney segmentation, Dice improved from $0.90\pm0.19$ to $0.95\pm0.10$ ($p < 4.01\times10^{-11}$). Multi-organ pooled performance showed a 68% variance reduction and a 74% improvement in boundary precision (Dice $0.95\pm0.07$). Crucially, frozen encoder models retained > 80% of fine-tuned performance without exposure to segmentation labels, proving the existence of learned anatomical priors. In low-data scenarios, diffusion-pretrained models maintained robust performance with only 50% (Dice: 0.92 liver, 0.94 kidney), 25%, and even 10% (Dice: 0.89 liver, 0.71 kidney) of labeled data. Using unlabeled images for diffusion-based pretraining successfully embeds robust anatomical features prior to human supervision, transforming U-Nets into anatomy-aware systems.
发表机构
- Manipal Institute of Technology Bengaluru, Manipal Academy of Higher Education(马尼帕尔理工学院班加罗尔分校,马尼帕尔高等教育学院)
机构由 AI 辅助整理,请以论文原文为准。