医学基础模型能否在非洲脑部数据上实现泛化?
Do Medical Foundation Models Generalize on the African Brain?
- Erasmus MC(伊拉斯姆斯大学医学中心)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
该研究评估医学基础模型在非洲脑部MRI数据上的泛化性,发现其无固有偏见,性能差异多源于数据集规模,核心是非洲神经影像数据集不足。
AI中文摘要:
医学基础模型(FMs)正越来越多地用于脑部MRI分析,但其评估仍以高资源数据集为主,针对非洲队列的泛化性研究不足。本研究在两项任务中评估FMs是否能同等泛化到非洲与非非洲脑部MRI数据:一是使用尼日利亚数据集进行痴呆分类,二是使用BraTS-Africa进行脑肿瘤分割。我们将两个通用型FMs(BrainIAC、3DINO)和两个分割专用FMs(MedSAM2、Medical-SAM2)与从头训练的基线模型对比。分类任务中,FMs带来的提升有限(BrainIAC的最高ROC-AUC为0.86);而分割任务中,FMs性能持续提升,MedSAM2的Dice系数最高达0.86。非洲与非非洲队列间的性能差异并不一致,似乎更多与数据集规模而非数据来源相关。这些结果表明FMs本身不存在针对非洲队列的固有偏见,同时强调非洲神经影像数据集的有限可用性与多样性是实现稳健评估和部署的主要障碍。
英文摘要:
Medical foundation models (FMs) are increasingly used for brain MRI analysis. However, their evaluation remains dominated by high-resource datasets, leaving generalization to African cohorts underexplored. We assess whether FMs generalize equally to African and non-African brain MRI data across two tasks: dementia classification using a Nigerian dataset and brain tumor segmentation using BraTS-Africa. We evaluate two generalist FMs (BrainIAC, 3DINO) and two segmentation-specific FMs (MedSAM2, Medical-SAM2) against a from-scratch baseline. For classification, FMs provide limited gains (highest ROC-AUC of 0.86 with BrainIAC), whereas for segmentation they consistently improve performance, reaching up to 0.86 Dice with MedSAM2. Performance differences between African and non-African cohorts are inconsistent and appear more related to dataset size than data origin. These results suggest that FMs do not exhibit an inherent bias against African cohorts, and highlight the limited availability and diversity of African neuroimaging datasets as the main barrier to robust evaluation and deployment.