基于不确定性感知视觉Transformer的多族裔近视与非近视人群青光眼检测:一项多中心模型开发与验证研究
Detecting Glaucoma Across Multi-ethnic Myopic and Non-Myopic Populations Using an Uncertainty-Aware Vision Transformer: A Multicentre Model Development and Validation Study
- Singapore Eye Research Institute(新加坡眼科研究所)
- Singapore National Eye Centre(新加坡国家眼科中心)
- Institute of Advanced Intelligence and Computing, Agency for Science, Technology and Research(新加坡科技研究局高级智能与计算研究院)
- Yong Loo Lin School of Medicine, National University of Singapore(新加坡国立大学杨潞龄医学院)
- Duke-NUS Medical School(杜克-新加坡国立大学医学院)
- Kaohsiung Chang Gung Memorial Hospital(高雄长庚纪念医院)
- Rothschild Foundation Hospital(罗斯柴尔德基金会医院)
- Beijing Visual Science and Translational Eye Research Institute (BERI)(北京视觉科学与转化眼科研究所)
- Beijing Tsinghua Changgung Hospital(北京清华长庚医院)
- Tsinghua University(清华大学)
- Suraj Eye Institute(苏拉杰眼科研究所)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本研究开发并验证了一种基于Vision Transformer的不确定性感知深度学习模型,利用彩色眼底照片在多种族近视与非近视人群中实现稳健的青光眼检测,其性能优于眼科医生,并有望支持高近视环境下的AI辅助筛查。
AI中文摘要:
背景:基于人工智能(AI)的彩色眼底照片(CFP)青光眼检测可提供可扩展的筛查,但由于真实标签定义、人群以及高度近视(HM)等共存疾病的差异,其性能在外部数据集上可能下降。我们开发并验证了一种基于Vision Transformer的深度学习(DL)模型,用于在有无高度近视的多族裔队列中进行青光眼检测。方法:使用56,483张CFP(57.1%为近视,14.4%为高度近视)开发了具有预测不确定性估计的ViT-B/16模型。青光眼标签通过临床、影像和视野检查数据进行标准化。该模型在三大洲的16个独立数据集上进行了验证,其中4个数据集具有明确的高度近视标签。结果:内部AUROC为98.7%(95% CI 98.2-99.1%),敏感性为94.5%,特异性为97.3%。在来自8个国家的16个外部数据集中,AUROC范围为86.4%至99.6%。在高度近视眼中,内部AUROC为97.8%(95% CI 96.1-99.2%),敏感性为94.8%,特异性为93.7%。外部高度近视AUROC在北京眼病研究中为86.5%,在台湾、泰国和韩国基于医院的数据集中分别为93.3%、91.8%和85.5%。在一项探索性高度近视临床评估中,该模型仅使用CFP的诊断准确性高于眼科医生和训练有素的评分员(92.0% vs 70.0%;p=0.008),且与使用完整临床信息的青光眼专家表现相当。解读:该模型在近视和非近视多族裔人群中显示出稳健的青光眼检测能力,并可能在高近视患病率环境中支持AI辅助筛查。
英文摘要:
Background: Artificial intelligence (AI)-based glaucoma detection from colour fundus photographs (CFP) offers scalable screening, but performance may decline on external datasets because of differences in ground-truth definitions, populations, and coexisting conditions such as high myopia (HM). We developed and validated a Vision Transformer-based deep learning (DL) model for glaucoma detection across multi-ethnic cohorts with and without HM. Methods: A ViT-B/16 model with predictive uncertainty estimation was developed using 56,483 CFPs (57.1% with myopia; 14.4% with HM). Glaucoma labels were standardised using clinical, imaging, and perimetry data. The model was validated on 16 independent datasets across three continents, including four datasets with explicit HM labels. Findings: Internal AUROC was 98.7% (95% CI 98.2-99.1%), with sensitivity 94.5% and specificity 97.3%. Across 16 external datasets from eight countries, AUROCs ranged from 86.4% to 99.6%. In HM eyes, internal AUROC was 97.8% (95% CI 96.1-99.2%), with sensitivity 94.8% and specificity 93.7%. External HM AUROCs were 86.5% in the Beijing Eye Study and 93.3%, 91.8%, and 85.5% in hospital-based datasets from Taiwan, Thailand, and South Korea. In an exploratory HM clinical evaluation, the model had higher CFP-only diagnostic accuracy than ophthalmologists and trained graders (92.0% vs 70.0%; p=0.008) and performed comparably to glaucoma specialists using full clinical information. Interpretation: The model showed robust glaucoma detection across myopic and non-myopic multi-ethnic populations and may support AI-assisted screening in settings with high HM prevalence.