arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.21300cs.CV

当适应适得其反:将表征漂移与MedSAM微调中的分布外故障关联起来

When Adaptation Hurts: Connecting Representational Drift to OOD Failures in MedSAM Fine-Tuning

  • University of Zagreb(萨格勒布大学)
  • University of Twente(特温特大学)

机构由 AI 辅助整理,请以论文原文为准。

Marko Haralović, Sounic Akkaraju, Carlo Baretta, Vasil Zapryanov, Alexia Briassouli

AI总结:

本研究分析MedSAM的六种微调策略在不同医学图像数据集上的泛化表现,发现仅编码器LoRA等策略可缓解远分布外性能下降,提示随机抖动能提升模型鲁棒性。

AI中文摘要:

用于医学图像分割的基础模型,如基于提示的MedSAM,在跨域和跨模态场景中表现出良好的泛化能力,通常适用于零样本或少样本设置。然而,它们的性能取决于提示的质量以及模型对自定义数据集的适应情况。本研究系统地考察了MedSAM在不同医学图像基准上的泛化能力,涉及六种适应策略:全模型LoRA、仅编码器LoRA、浅层视觉提示调优(VPT)、深层视觉提示调优(VPT)、仅解码器全微调、全模型微调。模型在国际皮肤成像协作挑战(ISIC 2018)数据集上进行训练,并在分布内(IN)和分布外(OOD)数据集上,在干净提示和噪声程度递增的提示下进行评估:近分布外(close-OOD)PH2(皮肤镜图像)、远分布外(far-OOD)BUSI(乳腺超声图像数据集)和CBIS-DDSM(数字化筛查乳腺图像数据库的精选乳腺成像子集)。我们发现,适应提升了分布内和近分布外数据的性能,但往往会降低远分布外数据的性能。全微调提供了最佳权衡,而仅编码器LoRA是最强的参数高效替代方案,在远分布外分布偏移下的表现优于标准LoRA和VPT。通过中心化核对齐(CKA),我们证明远分布外性能下降与解码器表征的漂移密切相关,而仅编码器的相似性无法解释鲁棒性。这表明仅编码器LoRA通过使编码器适应视觉特征的分布偏移,同时保留解码器通路,从而比标准LoRA提供更强的鲁棒性。我们进一步证明,对提示进行随机0-100像素抖动可产生更鲁棒、性能更优的模型。因此,我们得出结论,稳健的MedSAM适应需要综合考虑提示噪声暴露、域偏移和表征保存。我们发布了代码:this https URL

英文摘要:

Foundation models for medical image segmentation, like prompt-based MedSAM, generalize well across domains and modalities, often in zero or few-shot setups. However, their performance depends on the quality of prompts and the adaptation of the models to custom datasets. This work systematically examines how MedSAM generalizes across diverse medical imaging benchmarks, with six adaptation strategies: full-model and encoder-only LoRA, shallow and deep visual prompt tuning (VPT), and decoder-only and full fine-tuning. Models are trained on the International Skin Imaging Collaboration Challenge (ISIC 2018) dataset and evaluated under clean and increasingly noisy prompts on IN and Out-of-Distribution (OOD) datasets: close-OOD PH2 (dermoscopy), far-OOD BUSI (Breast Ultrasound Images Dataset) and CBIS-DDSM (Curated Breast Imaging Subset of the Digital Database for Screening Mammography). We show that adaptation improves performance on IN and close-OOD data but often reduces performance on far-OOD data. Full fine-tuning provides the best tradeoff, while encoder-only LoRA is the strongest parameter-efficient alternative, outperforming standard LoRA and VPT under far-OOD shifts. Using Centered Kernel Alignment (CKA), we show that far-OOD degradation is strongly associated with drift in decoder representations, whereas encoder similarity alone does not explain robustness. This suggests encoder-only LoRA provides stronger robustness than standard LoRA by adapting the encoder to distribution shift in visual features, while preserving the decoder pathway. We further show that random 0-100 pixel jitter on prompts produces more robust and better performing models. We thus conclude that robust MedSAM adaptation requires the combined consideration of prompt noise exposure, domain shift, and representation preservation. We release our code: https://github.com/ImSounic/medsam-vpt

补充信息

↑