arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

自我监督比临床监督更能推动医学基础模型中的表征趋同

Medical foundation models converge less under label supervision

Soroosh Tayebi Arasteh, Sebastian Ziegelmayer, Mahshad Lotfinia, Lisa Adams, Sven Nebelung, Jakob Nikolas Kather, Daniel Truhn

arXiv 2607.20274首次发表:更新:

发表机构

RWTH Aachen University; University Hospital RWTH Aachen; Technical University of Munich; Friedrich-Alexander-Universität Erlangen-Nürnberg; Technical University Dresden; University Hospital Dresden; University Hospital Heidelberg(亚琛工业大学; 亚琛工业大学附属医院; 慕尼黑工业大学; 埃尔朗根-纽伦堡弗里德里希-亚历山大大学; 德累斯顿工业大学; 德累斯顿大学附属医院; 海德堡大学附属医院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究探讨医学基础模型中表征趋同问题,通过对多个编码器剖析发现自我监督比临床监督更能驱动收敛,虽收敛有限但线性分类器可跨编码器转移,表明医学成像收敛由预训练目标决定,为互操作性设计与验证提供依据。

AI 中文摘要

不同团队的医学图像编码器越来越被视为可互换的,人们认为规模和临床监督会将其表征集中到一个共享结构上。但这种趋同是否真实、其产生原因以及是否具有临床可用性都未经检验,且相关相似性度量很脆弱。我们对18个图像和7个文本编码器进行了可控剖析,涵盖多种参数规模和成像模态。结果表明,收敛适度但高于随机水平,是由自我监督目标驱动,而非临床监督。线性分类器可跨编码器转移并应用于五家医院,保留约85%的编码器内性能。因此,医学成像中的收敛由预训练目标决定,而非规模或临床监督。应通过该目标设计互操作性,并在共享几何结构最薄弱的地方进行验证。

英文摘要

Diagnostic classifiers and imaging biomarkers are fitted on the embeddings of medical foundation models. These models are replaced as new versions appear. This practice assumes that different models represent images alike. We tested this assumption with more than 750,000 images from 14 datasets in five imaging modalities, 18 public models, and 101 models trained on chest radiographs and histopathology that differ in pretraining objective, label type, size, random seed, initialization, or training patients. Agreement was measured as the mutual k-nearest-neighbor overlap, the share of an image's nearest neighbors common to two models. Two public models shared on average 0.091 of the 10 nearest neighbors of a chest radiograph. A randomly initialized network shared 0.036 with them. In the four other modalities, they shared at most 0.300. Convergence depended on the pretraining objective. Models trained with label supervision converged least in all five controlled settings. On chest radiographs, two label-supervised models that differed only in their random seed shared 0.062 to 0.100 of their nearest neighbors. Two models trained with self-distillation, masked image modeling, or contrastive pretraining shared 0.411 to 0.802. On chest radiographs, agreement depended more on the source dataset and the radiographic view than on the findings. It was not significantly associated with accuracy. A linear mapping between the embeddings of two models, fitted on 4,096 unlabeled radiographs, transferred classifiers for 14 chest findings with 0.987 of their original area under the receiver operating characteristic curve. Convergence can therefore be chosen when a model is trained. A chest radiograph classifier can be transferred to a new model without new labels.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑