FOCUS:从基础视觉编码器到多模态大语言模型的视网膜模型泛化基准测试
FOCUS: Benchmarking Retinal Model Generalization from Foundation Vision Encoders to Multimodal LLMs
AI总结:
提出FOCUS跨数据集基准,系统评估视网膜眼底模型在数据集偏移下的泛化、校准与公平性,发现无单一模型家族全面占优,通用编码器平均性能最佳,微调提升域内但迁移不均。
AI中文摘要:
基于人工智能的视网膜图像分析随着基础模型的发展而取得进步,然而评估其可靠性仍然具有挑战性。在单一数据集上报告的性能并不能反映模型在数据集偏移、不同临床定义或不同患者亚组下的行为。这一局限性在医学影像分析中尤为关键,因为鲁棒性、校准性和公平性对于安全部署至关重要。我们提出了FOCUS(Foundation Ophthalmic Cross-Dataset Understanding under Shift,即偏移下基础眼科跨数据集理解),这是一个跨数据集基准测试,用于评估视网膜眼底模型,涵盖仅视觉编码器模型(VM)、视觉语言双编码器模型(VLM)和多模态大语言模型(MLLM)。FOCUS统一了二进制糖尿病视网膜病变、可转诊糖尿病视网膜病变和青光眼视神经病变任务,跨越十个公共数据集,这些数据集涵盖不同的地理区域、采集条件和标签协议。该基准测试通过一个统一分析层评估模型,该分析层衡量排序性能、校准性、亚组差异和图像质量鲁棒性。我们展示了一项大规模评估,涵盖532个基础配置和228个通过低秩适配(LoRA)进行监督微调的MLLM配置。结果表明,没有模型家族在任务和数据集中持续占主导地位:通用VM编码器实现了最强的平均排序性能,医学MLLM具有竞争力但表现可变,而双编码器VLM则从轻量级适配中显著受益。微调提高了域内性能,但在外部数据集上表现出异质迁移,尤其是在校准性方面。这些发现表明视网膜模型评估本质上是多维的。FOCUS提供了一个实用框架和公共基准,用于评估超越单一数据集排行榜的泛化性、可靠性和鲁棒性。
英文摘要:
Progress in AI-based retinal image analysis has advanced with foundation models, yet evaluating their reliability remains challenging. Performance reported on a single dataset does not capture how models behave under dataset shift, across clinical definitions, or for different patient subgroups. This limitation is particularly critical in medical imaging analysis, where robustness, calibration, and fairness are essential for safe deployment. We introduce FOCUS (Foundation Ophthalmic Cross-Dataset Understanding under Shift), a cross-dataset benchmark for evaluating retinal fundus models that considers vision-only encoder models (VM), vision-language dual-encoder models (VLM), and multimodal large language models (MLLM). FOCUS harmonizes binary diabetic retinopathy, referable diabetic retinopathy, and glaucomatous optic neuropathy tasks across ten public datasets spanning diverse geographies, acquisition conditions, and label protocols. The benchmark evaluates models through a unified analysis layer that measures ranking performance, calibration, subgroup disparities, and image-quality robustness. We present a large-scale evaluation covering 532 base configurations and 228 MLLM configurations adapted through supervised fine-tuning with low-rank adaptation (LoRA). Results show that no model family consistently dominates across tasks and datasets: general VM encoders achieve the strongest average ranking performance, medical MLLMs are competitive but variable, and dual encoder VLMs benefit substantially from lightweight adaptation. Fine-tuning improves in-domain performance but exhibits heterogeneous transfer to external datasets, particularly in calibration. These findings demonstrate that retinal model evaluation is inherently multidimensional. FOCUS provides a practical framework and public benchmark to assess generalization, reliability, and robustness beyond single-dataset leaderboards