arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.15195cs.CV

超越自然图像基础模型:针对眼科图像分析的卫星图像预训练基准测试

Beyond Natural-Image Foundation Models: Benchmarking Satellite Pretraining for Ophthalmic Image Analysis

  • Faculty of Electrical Engineering and Computing, University of Zagreb(萨格勒布大学电气工程与计算机学院)
  • University College London(伦敦大学学院)
  • Institute of Ophthalmology, University College London(伦敦大学学院眼科研究所)
  • Department of Computer Science, University College London(伦敦大学学院计算机科学系)
  • NIHR Moorfields Biomedical Research Centre(NIHR穆尔菲尔德生物医学研究中心)
  • Hawkes Institute, University College London(伦敦大学学院霍克斯研究所)
  • Moorfields Eye Hospital NHS Foundation Trust(穆尔菲尔德眼科医院NHS基金会信托)

机构由 AI 辅助整理,请以论文原文为准。

Lovre Antonio Budimir, Mingya Alexa Gong, Alyssa Foong Quinney, Ivana Matovinović, Yukun Zhou, Pearse A. Keane, Sven Lončarić, Marinko V. Šarunić

AI总结:

本文针对眼科图像分析,对比卫星图像与自然图像预训练的视觉基础模型,发现卫星图像预训练在眼科任务上表现优于自然图像,部分任务可媲美医学专家模型。

AI中文摘要:

视觉基础模型(VFMs)已成为医学成像领域颇具前景的方法,可生成适用于各类成像模态、解剖区域及临床任务的通用系统,且能高效适配不同场景。不过,VFMs需要海量训练数据,而医学图像分析领域受限于数据可得性、隐私问题及高昂开发成本,其发展受到制约。为缓解这些限制,医学视觉基础模型(MedVFMs)通常基于在大量公开自然图像上预训练的通用模型权重构建,这会给医学任务适配带来显著的分布偏移问题。针对这一问题,本文提出将卫星图像作为开发和基准测试MedVFM的新型预训练领域,其动机在于卫星图像与医学数据的视觉契合度更高,且不受医学数据集存在的隐私限制。在多种眼科成像模态上,本文对比了在4.93亿张卫星图像上预训练的DINOv3-SAT493m,与在17亿张自然图像上预训练的DINOv3-LVD1689m,同时还对比了两个医学专家基线模型:DINOv3-RETFound和MAE-RETFound。实验结果表明,对于眼科任务而言,卫星图像作为预训练源的表现优于自然图像,尤其在富含血管的正面成像模态上;在多项任务中,即便未使用任何医学数据,卫星图像预训练在高分辨率正面输入上的表现可与医学专家模型相媲美甚至超出。

英文摘要:

Vision Foundation Models (VFMs) have emerged as a promising approach in medical imaging, producing broadly applicable systems that can be efficiently adapted across diverse imaging modalities, anatomical regions, and clinical tasks. However, VFMs require extensive training data, and their progress in medical image analysis is constrained by limited data availability, privacy concerns, and high development costs. To alleviate these constraints, medical VFMs (MedVFMs) are often built upon weights from generalist models pretrained on vast amounts of publicly available natural images, introducing a substantial distribution shift for medical task adaptation. To address this, we propose satellite imagery as a novel pretraining domain for MedVFM development and benchmarking, motivated by its closer visual alignment with medical data and its freedom from the privacy constraints that limit medical datasets. Across multiple ophthalmic imaging modalities, we compare DINOv3-SAT493m pretrained on 493 million satellite images against DINOv3-LVD1689m pretrained on 1.7 billion natural images, together with two medical specialist baselines: DINOv3-RETFound and MAE-RETFound. Our experiments show that satellite imagery is a stronger pretraining source than natural images for ophthalmic tasks, particularly on en face vascular-rich modalities. On several tasks, satellite pretraining matches or exceeds the medical specialists on high-resolution en face inputs, despite using no medical data.

补充信息

↑