arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

在域转移下对用于乳腺钼靶成像的基础模型的稳健性进行基准测试

Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift

Giang Nguyen, Raghav Mehta, Emma A. M. Stanley, Tian Xia, Thi Hao Nguyen, Hieu Pham, Ben Glocker

arXiv 2607.10358首次发表:更新:

发表机构

College of Engineering and Computer Science, VinUniversity; Imperial College London; Radiology Department, Vietnam National Cancer Hospital; VinUni-Illinois Smart Health Center, VinUniversity; The Computer Vision and Medical AI Lab, VinUniversity(工程与计算机科学学院,文大大学; 伦敦帝国理工学院; 越南国家癌症医院放射科; 文大大学 - 伊利诺伊智能健康中心,文大大学; 计算机视觉与医学人工智能实验室,文大大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究在域转移下乳腺钼靶成像基础模型的稳健性,用统一协议在多数据集上训练评估15种模型主干,发现特定视觉语言模型性能强,DINOv3是有竞争力基线,适应预训练未持续提升泛化,强调数据集级OOD评估是核心标准。

AI 中文摘要

基础模型越来越多地被用作乳腺钼靶成像的图像特征提取器,但其在外部域转移下的稳健性仍不明确。我们使用统一的冻结主干线性探针协议,在3个源数据集上训练,并在标签协调后在12个任务兼容的分布外(OOD)数据集上评估,对15种基础模型主干在乳腺密度、BI-RADS严重程度和癌症状态方面进行基准测试。乳腺钼靶特定的视觉语言模型(Mammo-FM和MaMA)提供了最强的平均OOD性能,但稳健性不能仅由乳腺钼靶暴露来解释。DINOv3仍然是一个有竞争力的仅视觉基线,并且适应乳腺钼靶的预训练并不能持续提高泛化能力。数据集级分析进一步表明,即使是领先模型在不同数据集上的性能也存在异质性。特征空间检查表明,有用的表示可以在保留数据集和采集结构的同时保留临床信号。这些发现突出了数据集级OOD评估作为评估乳腺钼靶表示的核心标准。我们的代码可公开获取:此https URL。

英文摘要

Foundation models are increasingly used as image feature extractors for mammography, but their robustness under external domain shift remains unclear. We benchmark 15 foundation-model backbones across breast density, BI-RADS severity, and cancer status using a unified frozen-backbone linear-probe protocol, training on 3 source datasets and evaluating on 12 task-compatible out-of-distribution (OOD) datasets after label harmonization. Mammography-specific vision-language models (Mammo-FM and MaMA) provide the strongest mean OOD performance, but robustness is not explained by mammography exposure alone. DINOv3 remains a competitive vision-only baseline, and mammography-adapted pretraining does not consistently improve generalization. Dataset-level analysis further shows that even leading models show heterogeneous performance across datasets. Feature-space inspection reveals that useful representations can preserve clinical signal while retaining dataset and acquisition structure. These findings highlight dataset-level OOD evaluation as a central criterion for assessing mammography representations. Our code is publicly available: https://github.com/biomedia-mira/mammo-ood.

CommentsAccepted at Deep-Brea3th 2026 workshop in conjunction with MICCAI 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑