arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.25810cs.CV

分布偏移下基于伪标签差异的医学图像分类无基准模型选择

Label-Free Foundational Model Selection for Medical Image Classification under Distribution Shift via Pseudo Label Discrepancy

Juan Iñaki Larrea, Lucas Mansilla, Enzo Ferrante

首次发表
浏览论文内容

中文总结 AI 辅助

针对分布偏移下医学图像分类的无基准模型选择问题,提出基于SUDO框架的AURCC准则,在胸部X射线分类任务中实现了与真实排名高相关性的模型排名,小源标注数据下表现更优。

中文摘要 AI 辅助

基础模型越来越多地被部署用于医学图像分析,但在部署时典型的机构间分布偏移场景下,它们的性能差异很大,且在目标域标签极少可用的情况下无法得知其性能,这留下了一个未解决的实际问题:给定多个候选基础模型和源域的标注数据,应选择哪一个部署到无标注的目标域?我们提出一种无基准选择准则,该准则基于SUDO框架(一种无需真实标注即可评估临床AI系统的框架)。SUDO通过预测概率划分无标注的目标数据,并针对每个区域测量反映类别污染的伪标签差异;跨区域聚合后得到一个既不需要目标标注也不需要微调的分数(AURCC)。我们在三种跨医院偏移场景下的胸部X射线分类任务中,于零样本和MLP探测 regime下,展示了AURCC可用于对多种视觉语言模型(BioMedCLIP、CXR-CLIP、CheXzero、MedCLIP、MedImageInsight、CLIP)进行排名。AURCC排名与真实排名的Spearman相关系数最高达0.943(p<0.05)。与通过保留源域准确率进行排名的自然基线相比,当标注源域数据量较大时AURCC具有竞争力,而当标注源域数据量较小时,AURCC能产生更准确的排名——这是资源受限场景下的关注 regime。

英文摘要

Foundation models are increasingly deployed for medical image analysis. However, under the inter-institutional distribution shift typical of deployment, their performance varies widely and cannot be known without target-domain labels, which are rarely available. This leaves a practical question unresolved: given several candidate foundational models and labeled-data from a source domain, which one to deploy in an unlabeled target domain? We propose a label-free selection criterion built on SUDO, a framework for evaluating clinical AI systems without ground-truth annotations. SUDO partitions the unlabeled target data by predicted probability and, for each region, measures a pseudo-label discrepancy reflecting class contamination; aggregated across regions, this yields a score (AURCC) requiring neither target annotation nor fine-tuning. We show that AURCC can be used to rank a variety of vision-language models (BioMedCLIP, CXR-CLIP, CheXzero, MedCLIP, MedImageInsight, CLIP) on chest X-ray classification across three inter-hospital shift scenarios, under zero-shot and MLP-probe regimes. The AURCC ranking recovers the ground-truth ranking with Spearman rho up to 0.943 (p<0.05). Against the natural baseline of ranking by held-out source accuracy, AURCC is competitive when the labeled source is large and yields a more accurate ranking once it is small; the regime of interest in resource-constrained settings.

发表机构

  • Universidad de Buenos Aires(布宜诺斯艾利斯大学)
  • CONICET(阿根廷国家科学技术研究委员会)
  • Universidad Nacional del Litoral(国立 littoral 大学)
  • Research Institute for Signals, Systems and Computational Intelligence(信号、系统与计算智能研究所)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑