ORViT-DR:用于低分辨率糖尿病视网膜病变分级的序数鲁棒混合视觉Transformer模型
ORViT-DR: Ordinally-Robust Hybrid ViT for Low-Resolution Diabetic Retinopathy Grading
浏览论文内容
中文总结 AI 辅助
本研究提出ORViT-DR混合深度学习框架,结合卷积与Transformer技术,在RetinaMNIST数据集上实现低分辨率DR分级,取得57.00%准确率等结果,验证了混合架构的有效性。
中文摘要 AI 辅助
糖尿病视网膜病变(DR)是视力受损的主要原因之一,一套可靠的自动分级系统可使筛查过程更安全、准确。由于DR分期呈渐进式发展,疾病严重程度分级任务天然具有序数结构,相邻类别视觉特征相似。本研究设计了混合深度学习框架ORViT-DR,用于从低分辨率视网膜图像中提升DR分级性能。该方法通过预训练的ViT-Hybrid骨干网络,结合卷积特征提取与基于Transformer的全局上下文建模,该骨干网络集成了BiT-ResNetv2与Vision Transformer架构。模型在MedMNISTv2数据集的RetinaMNIST子集上进行测试,该子集包含28×28的视网膜眼底图像,标注了5个疾病严重程度等级。为促进稳定训练与更好的特征学习,训练策略采用渐进式层解冻、分层学习率衰减、指数移动平均(EMA)参数更新及推理时的集成预测。在官方RetinaMNIST测试集上的实验结果显示,该方法的分类准确率达57.00%,二次加权Kappa评分为0.5963,宏F1评分为0.4293。这些结果表明,混合CNN-Transformer架构可为序数视网膜图像分析提供有效表示。
英文摘要
Diabetic retinopathy (DR) is one of the main causes of impaired vision. A good and reliable automated grading system can make the screening process safer and more accurate. Because DR stages progress gradually, the task of grading disease severity naturally follows an ordinal structure in which neighboring classes share similar visual characteristics. In this study, ORViT-DR, a hybrid deep learning framework, is designed to improve DR grading from low-resolution retinal images. The proposed approach combines convolutional feature extraction with transformer-based global context modeling through a pre-trained ViT-Hybrid backbone, which integrates BiT-ResNetv2 with a Vision Transformer architecture. The approach is tested on the RetinaMNIST subset of the MedMNISTv2 dataset, which contains 28x28 retinal fundus images annotated with five levels of disease severity. To promote stable training and better feature learning, the training strategy applies progressive layer unfreezing, layer-wise learning rate decay, exponential moving average (EMA) parameter updates, and ensemble-based prediction during inference. Experimental results on the official RetinaMNIST test set show that the proposed method achieves 57.00% classification accuracy, along with a quadratic weighted kappa score of 0.5963 and a macro-F1 score of 0.4293. These results suggest that hybrid CNN-Transformer architectures can provide effective representations for ordinal retinal image analysis.
发表机构
- Dhaka International University(达卡国际大学)
- KUET(库尔纳工程技术大学)
- Bangladesh Bank(孟加拉银行)
机构由 AI 辅助整理,请以论文原文为准。