arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

用于超广域视网膜成像的基础模型表示迁移

Representation Transfer of Foundation Models for Ultra-Widefield Retinal Imaging

Mingya Alexa Gong, Da Ma, Lovre Antonio Budimir, Ivana Matovinovic, Sven Loncaric, Myeong Jin Ju, Yukun Zhou, Siegfried K. Wagner, Pearse A. Keane, Marinko V. Sarunic

arXiv 2608.00586首次发表:更新:

发表机构

Institute of Ophthalmology, University College London; Wake Forest University School of Medicine; Virginia Tech-Wake Forest University School of Biomedical Engineering and Sciences; University of Zagreb Faculty of Electrical Engineering and Computing; NIHR Biomedical Research Centre, Moorfields Eye Hospital NHS Foundation Trust(伦敦大学学院眼科研究所; 维克森林大学医学院; 弗吉尼亚理工大学-维克森林大学生物医学工程与科学学院; 萨格勒布大学电气工程与计算机学院; NIHR生物医学研究中心,摩尔菲尔兹眼科医院NHS基金会信托)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究探讨基础模型预训练策略对超广域视网膜成像表示迁移的影响,发现DINOv3模型在五类糖尿病视网膜病变分级中性能最优,监督和自蒸馏预训练的ViT优于MAE,部分微调可缩小MAE的性能差距。

AI 中文摘要

尽管基础模型作为医学成像的特征提取器已被广泛应用,但人们对不同预训练策略如何影响学习到的表示向弱监督眼科成像任务的迁移性却知之甚少。我们在超广域(UWF)视网膜成像中研究了这一问题,在基于补丁的多实例学习(MIL)框架内评估基础模型表示,用于对UWF图像进行疾病分类。我们比较了以监督、掩码自编码器(MAE)和自蒸馏目标预训练的视觉Transformer(ViT)编码器,同时保持下游聚合架构不变。在对以ImageNet-1k预训练的ViT-B编码器的受控比较中,预训练目标的选择显著影响了冻结表示的迁移,其中基于监督和自蒸馏的模型表现优于MAE。以更大规模预训练的当代DINOv3模型实现了整体最强性能,其五类糖尿病视网膜病变分级的二次加权kappa为0.863,与DINOv1相当。注意力分析进一步揭示了与不同预训练表示相关的独特补丁聚合行为,而部分微调显著缩小了MAE的性能差距。这些发现表明,预训练策略既影响表示迁移性,也影响MIL中补丁级证据的后续聚合,进而导致下游分类性能的差异。

英文摘要

Despite the widespread adoption of foundation models as feature extractors for medical imaging, relatively little is understood about how different pretraining strategies influence the transferability of learned representations to weakly supervised ophthalmic imaging tasks. We investigate this question in ultra-widefield (UWF) retinal imaging by evaluating foundation model representations within a patch-based multiple instance learning (MIL) framework for disease classification on UWF images. We compare Vision Transformer encoders pretrained with supervised, Masked Autoencoder (MAE), and self-distillation objectives, while keeping the downstream aggregation architecture unchanged. Within a controlled comparison of ViT-B encoders pretrained on ImageNet-1k, the choice of pretraining objective substantially influenced frozen representation transfer, with supervised and self-distillation-based models outperforming MAE. A contemporary DINOv3 model pretrained at a larger scale achieved the strongest overall performance, with a quadratic weighted kappa of 0.863 for five-class diabetic retinopathy grading, comparable with DINOv1. Attention analysis further revealed distinct patch-aggregation behaviours associated with the different pretrained representations, while partial fine-tuning substantially reduced the performance gap for MAE. These findings suggest that pretraining strategy influences both representation transferability and the subsequent aggregation of patch-level evidence within MIL, resulting in differences in downstream classification performance.

Comments15 pages, 7 figures

Journal refM. A. Gong, D. Ma, L. A. Budimir, I. Matovinovic, S. Loncaric, M. J. Ju, Y. Zhou, S. K. Wagner, P. A. Keane, and M. V. Sarunic, "Representation Transfer of Foundation Models for Ultra-Widefield Retinal Imaging," IEEE Access, 2026

DOI:10.1109/ACCESS.2026.3737380

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑