发表机构
Department of Computer Science, Memorial University of Newfoundland; Division of Biomedical Sciences, Faculty of Medicine, Memorial University of Newfoundland(纽芬兰纪念大学计算机科学系; 纽芬兰纪念大学医学院生物医学科学部)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究乳腺钼靶筛查人工智能中跨数据集情况,用EfficientNet - B5编码器等方法评估,发现添加外部阳性病例会降性能,重新定义任务发现数据集特征影响表示,表明合并数据集有领域偏移,需领域感知策略。
AI 中文摘要
可靠的乳腺钼靶筛查人工智能需要具有低癌症患病率和筛查人群中细微异常特征的训练数据。我们研究了用来自异常丰富的外部数据集的活检确诊病例补充此类数据是否能提高性能。我们使用纽芬兰和拉布拉多乳腺筛查数据集(NLBSD)以及CBIS - DDSM和CMMD,评估了以Mammo - CLIP权重初始化的EfficientNet - B5编码器作为固定线性探测器,采用一致的预处理和患者级分割。仅使用NLBSD的模型AUC - ROC为0.737。添加外部阳性病例会降低性能,领域匹配评估仅在训练和测试领域一致时才有适度提升,且无配置超过仅使用NLBSD的模型。将任务重新定义为预测每个检查的数据集来源,发现数据集特定特征强烈影响学习表示。这表明单纯合并异常丰富的乳腺钼靶数据集会引入领域偏移,归一化后采集、强度映射和数据集构建的差异依然存在,因此需要领域感知策略来组合异构乳腺钼靶数据集。
英文摘要
Reliable AI for screening mammography requires training data representative of the low cancer prevalence and subtle abnormalities found in screening populations. We examined whether supplementing such data with biopsy-confirmed cases from abnormal-enriched external datasets improves performance. Using the Newfoundland and Labrador Breast Screening Dataset (NLBSD) alongside CBIS-DDSM and CMMD, we evaluated an EfficientNet-B5 encoder initialized with Mammo-CLIP weights as a frozen linear probe under consistent preprocessing and patient-level splits. The NLBSD-only model achieved an AUC-ROC of 0.737 (95% CI [0.686, 0.785]). Adding external positive cases reduced performance in every configuration (AUC-ROC = 0.620--0.644; DeLong test, Holm-corrected $p < 0.05$), with degradation increasing as additional sources were introduced. Domain-matched evaluation produced modest gains only when the training and test domains coincided, and no configuration surpassed the NLBSD-only model. As a diagnostic, we reframed the task as predicting each examination's dataset of origin. The datasets were separated almost perfectly despite identical preprocessing, indicating that dataset-specific characteristics strongly influence the learned representation. These findings show that naïvely pooling abnormal-enriched mammography datasets can introduce domain shift that outweighs the benefit of additional positive cases. Differences in acquisition, intensity mapping, and dataset construction persist after normalization, motivating domain-aware strategies for combining heterogeneous mammography datasets.