arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

医学影像中迁移性评估指标的鲁棒性

Robustness of transferability estimation metrics for medical imaging

Niclas Claßen, Théo Sourget, Dovile Juodelyte, Rob van der Goot, Veronika Cheplygina

arXiv 2608.09999首次发表:更新:

发表机构

IT University of Copenhagen(哥本哈根信息技术大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对医学影像迁移学习,探究目标数据集及评估指标对迁移性评估指标排序的影响,发现其与参考排序一致性低,揭示了TE指标鲁棒性的不足。

AI 中文摘要

在迁移学习中,源模型的选择在很大程度上影响目标数据集上的性能,然而选择合适的源模型仍是一项具有挑战性的任务,尤其在医学影像领域,需在ImageNet等现成选项预训练的模型与特定领域数据集预训练的模型之间做出抉择。迁移性评估(TE)指标旨在以计算成本效益高的方式预测性能最佳的源模型,从而解决这一问题。但此前研究报告的TE指标性能存在矛盾,原因在于实验设置存在差异。此外,大多数TE指标是为自然图像设计并在自然图像上评估的,且针对准确率优化,而医学影像中通常使用对类别不平衡更鲁棒的指标。本研究通过构建不同样本量和随机种子的微型总体,将目标数据集作为孤立因素研究其影响;同时,探究用于获取参考排序的评估指标的影响。研究发现,目标数据集的微小修改会改变排序,且评估指标的选择会影响参考排序,进而影响TE指标的评估。总体而言,TE指标排序与参考排序之间的一致性较低。本工作使用的代码、模型检查点和数据划分可通过此https URL获取。

英文摘要

In transfer learning, the choice of source model largely influences the performance on a target dataset. Still, selecting a fitting source remains a challenging task, especially in medical imaging where one has to decide between models pre-trained on off-the-shelf options, such as ImageNet, and domain specific datasets. Transferability estimation (TE) metrics address this problem by aiming to predict the best performing source model in a computationally cost effective way. However, previous work has reported conflicting TE metric performances due to differences in experimental setups. Moreover, most TE metrics are designed for and evaluated on natural images, while being optimized for accuracy, whereas in medical imaging metrics that are more robust to class imbalance are typically used. We study the impact of varying the target dataset as an isolated factor, by constructing miniature populations of different sample sizes and random seeds. In addition, we investigate the influence of the evaluation metric used to obtain the reference ranking. We find that small modifications to the target dataset change the rankings. Furthermore, we show that the choice of evaluation metric affects the reference rankings and therefore the evaluation of TE metrics. Overall, we observe a low agreement between rankings from TE metrics and reference. The code, model checkpoints and data splits used in this work are available through https://github.com/niclasclassen/robustness-of-transferability-estimation-metrics-for-medical-imaging.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑