三维医学基础模型能看穿MRI伪影吗?一项关于表示鲁棒性的对照研究
Do 3D Medical Foundation Models See Through MRI Artifacts? A Controlled Study of Representation Robustness
- Pioneer Centre for AI, University of Copenhagen(哥本哈根大学先锋人工智能中心)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文对照评估五种三维医学基础模型的表示鲁棒性,发现其鲁棒性与模型及伪影类型相关,部分模型对伪影更敏感,仅靠大规模预训练无法保证伪影不变性,需部署前评估鲁棒性。
AI中文摘要:
自监督三维医学基础模型正日益被用作通用特征提取器,但其对MRI伪影的敏感性仍未得到充分理解。本文对五种预训练三维编码器的表示鲁棒性开展对照评估,这些编码器涵盖不同架构、目标、预训练领域及数据集规模。研究使用带有四种MRI序列的BraTS-Africa病例,生成七种频率域和图像域伪影,设置五个预定义的损坏程度。通过线性中心核对齐(CKA)、RankMe和UMAP评估鲁棒性,辅以独立的分割一致性分析。研究发现,鲁棒性具有强烈的模型和伪影依赖性:3DINO表现出最一致稳定的表示,BrainIAC对多种损坏高度敏感,NeuroVFM、BrainFM和Neuro-SimCLR则呈现中等但不同的伪影特异性特征。在多数条件下,CKA大幅下降而RankMe相对稳定,表明伪影常扭曲表示几何结构但不会导致维度崩溃。分割一致性在损坏下也会降低,尤其在重影和Rician噪声下,但与表示层面鲁棒性仅部分一致。这些发现表明,仅靠更大规模或领域特定的预训练无法保证伪影不变性,且为三维基础模型部署到异构MRI环境前开展明确的鲁棒性评估提供了动机。
英文摘要:
Self-supervised 3D medical foundation models are increasingly used as general-purpose feature extractors, yet their sensitivity to MRI artifacts remains poorly understood. We present a controlled evaluation of representation robustness across five pretrained 3D encoders spanning different architectures, objectives, pretraining domains, and dataset scales. Using BraTS-Africa cases with four MRI sequences, we generate seven frequency- and image-domain artifacts at five predefined corruption settings. Robustness is assessed using linear centered kernel alignment (CKA), RankMe, and UMAP, complemented by an independent segmentation-consistency analysis. We find that robustness is strongly model- and artifact-dependent. 3DINO exhibits the most consistently stable representations, while BrainIAC is highly sensitive to several corruptions; NeuroVFM, BrainFM, and Neuro-SimCLR show intermediate but distinct artifact-specific profiles. Across many conditions, CKA decreases substantially while RankMe remains comparatively stable, indicating that artifacts often distort representation geometry without causing dimensional collapse. Segmentation consistency also degrades under corruption, particularly for ghosting and Rician noise, but aligns only partially with representation-level robustness. These findings show that larger-scale or domain-specific pretraining alone does not guarantee artifact invariance and motivate explicit robustness evaluation before deploying 3D foundation models in heterogeneous MRI settings.