基于轻量级视觉Transformer的U-Net用于MRI脑肿瘤分割
Lightweight Vision Transformer-Based U-Net for Brain Tumor Segmentation from MRI
- IUBAT(国际伊斯兰大学孟加拉国分校)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
提出一种轻量级ViT-UNet混合架构,以仅260万参数在MRI脑肿瘤分割中实现全局上下文建模,并在TCGA LGG数据集上以mIoU 0.8100和Dice 0.8446超越基线UNet。
AI中文摘要:
从磁共振成像中准确分割脑肿瘤对于诊断、治疗规划和手术指导至关重要。尽管卷积神经网络,尤其是UNet,在医学图像分割中取得了显著成功,但它们往往难以捕捉建模具有不规则形状和复杂边界的肿瘤所需的长程空间依赖性。本文提出了一种轻量级视觉Transformer UNet,将UNet的分层特征提取能力与视觉Transformer的全局上下文建模相结合。所提出的架构在U-Net编码器-解码器框架内集成了一个紧凑的ViT瓶颈,能够有效学习局部和全局特征,同时仅需260万个可训练参数即可保持计算效率。该模型在TCGA LGG MRI分割数据集上进行了评估,实现了平均交并比0.8100和Dice分数0.8446,分别比基线UNet高出3.75%和3.15%。广泛的定量和定性分析,包括混淆矩阵评估、精确率-召回率曲线、逐图像性能分布和肿瘤大小依赖性分析,证明了所提出方法在脑肿瘤分割中的有效性和鲁棒性。
英文摘要:
Accurate brain tumor segmentation from Magnetic Resonance Imaging is essential for diagnosis, treatment planning, and surgical guidance. Although Convolutional Neural Networks, particularly UNet, have achieved significant success in medical image segmentation, they often struggle to capture the long-range spatial dependencies required to model tumors with irregular shapes and complex boundaries. This paper proposes a lightweight Vision Transformer UNet that combines the hierarchical feature extraction capability of UNet with the global context modeling of Vision Transformers. The proposed architecture incorporates a compact ViT bottleneck within a U-Net encoder-decoder framework, enabling effective learning of both local and global features while maintaining computational efficiency with only 2.6 million trainable parameters. The model was evaluated on the TCGA LGG MRI Segmentation dataset, achieving a mean Intersection over Union of 0.8100 and a Dice score of 0.8446, outperforming the baseline UNet by 3.75% and 3.15%, respectively. Extensive quantitative and qualitative analyses, including confusion matrix evaluation, precision recall curves, per-image performance distribution, and tumor size dependency analysis, demonstrate the effectiveness and robustness of the proposed method for brain tumor segmentation.