arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

MCSeg:用于多模态心脏图像分割的体积金字塔Transformer的预训练与微调

MCSeg: Pre-training and Fine-tuning Volumetric Pyramid Transformer for Multi-modal Cardiac Image Segmentation

Zhiyu Ye, Hairong Zheng, Tong Zhang

arXiv 2608.30371首次发表:更新:

发表机构

Shenzhen Institute of Advanced Technology; Pengcheng Laboratory; University of Chinese Academy of Sciences(深圳先进技术研究院; 鹏城实验室; 中国科学院大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出MCSeg网络,通过新型SFP连接ViT编码器与CNN解码器,结合自监督预训练与RMI损失,在多模态心脏图像分割任务中优于11种SOTA方法,小样本场景也表现出色

AI 中文摘要

心脏图像自动分割对心脏疾病的诊断与治疗至关重要。本研究提出MCSeg,一种专为多模态心脏分割设计的基于体积Transformer的网络。为克服现有混合网络固有的架构不匹配问题,我们提出一种新型缩放特征金字塔(Scaling Feature Pyramid, SFP)。与传统跳跃连接不同,SFP通过将ViT的输出转换为分层特征金字塔,有效连接单尺度3D视觉Transformer(Vision Transformer, ViT)编码器与多尺度CNN解码器,确保全局上下文信息得到有效利用。在训练范式上,ViT编码器首先通过掩码图像建模进行自监督预训练,随后网络在下游任务上进行微调,此过程中整合区域互信息(Regional Mutual Information, RMI)损失以提升边界分割精度。实验中,MCSeg在CT数据集ImageCHD、多模态数据集MM-WHS、MRI数据集HVSMR-2.0及MSD Heart上均持续优于11种SOTA方法,凸显其在多模态心脏分割任务中的有效性;此外,MCSeg在小样本实验中的优异性能,展现出其在有限数据场景下的显著适应潜力。代码与预训练ViT-B权重已开源至该https URL

英文摘要

Automatic cardiac image segmentation is pivotal for diagnosing and treating cardiac diseases. In this work, we introduce MCSeg, a volumetric transformer-based network tailored for multi-modal cardiac segmentation. To overcome the architectural mismatch inherent in existing hybrid networks, we propose a novel Scaling Feature Pyramid (SFP). Unlike conventional skip connections, the SFP effectively bridges the single-scale 3D Vision Transformer (ViT) encoder and the multi-scale CNN decoder by transforming the ViT's output into a hierarchical feature pyramid, ensuring that global contextual information is effectively leveraged. For the training paradigm, the ViT encoder first undergoes self-supervised pre-training via masked image modeling. Subsequently, the network is fine-tuned on downstream tasks, during which a regional mutual information (RMI) loss is integrated to improve boundary segmentation accuracy. In experiments, MCSeg consistently outperforms eleven SOTA methods on CT dataset ImageCHD, multi-modal dataset MM-WHS, MRI dataset HVSMR-2.0 and MSD Heart, highlighting the effectiveness of our MCSeg for multi-modal cardiac segmentation tasks. Furthermore, MCSeg's superior performance in few-shot experiment showcases its significant potential in adapting to limited data scenarios. Codes and pre-trained ViT-B weights are open-sourced at https://openi.pcl.ac.cn/OpenMedIA/MCSeg

CommentsAccepted for publication in IEEE Journal of Biomedical and Health Informatics

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑