arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

注意力增强深度特征与异构集成学习用于青光眼检测

Attention-Enhanced Deep Features with Heterogeneous Ensemble Learning for Glaucoma Detection

Abdullah Al Shafi, Nishat Sadaf Lira, Abrar Hasan, Kazi Saeed Alam, Swapnil Kundu Argha

arXiv 2609.06699首次发表:更新:

发表机构

Khulna University of Engineering & Technology; Daffodil International University; Green University of Bangladesh(库尔纳工程技术大学; 水仙国际大学; 孟加拉国绿色大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出融合CBAM注意力增强深度特征与异构集成学习的青光眼检测框架,在公共眼底数据集上取得最优性能,并利用Grad-CAM提供可解释性证据。

AI 中文摘要

青光眼是一种进行性视神经病变,其特征是视神经的不可逆损伤,因此及时诊断对于防止永久性视力丧失至关重要。尽管深度学习在自动化青光眼检测中已展现出良好的性能,但现有方法往往忽视特征细化、遭受类别不平衡问题,并依赖单一分类器从而限制了预测的鲁棒性。为解决这些挑战,本文提出了一种混合青光眼检测框架,该框架将注意力增强的深度特征提取与异构集成学习相结合。具体而言,使用InceptionV3提取深度表示,随后通过引入卷积块注意力模块(CBAM)进行细化,以增强判别性视网膜特征。为提高分类鲁棒性,提取的特征使用多种机器学习模型以及单层集成(SLE)和双层集成(DLE)策略进行分类,同时采用SMOTE结合Tomek Links(SMOTE+TL)来缓解类别不平衡。此外,对手工特征、深度特征和注意力增强深度特征表示进行了系统性比较。在两个公共视网膜眼底数据集上的实验评估表明,基于深度特征的方法始终优于基于手工特征的方法,而所提出的注意力增强框架实现了最佳的整体性能。此外,Grad-CAM可视化证实了所提出的模型聚焦于临床相关的视网膜区域,为模型的预测过程提供了可解释的证据。

英文摘要

Glaucoma is a progressive optic neuropathy characterized by irreversible damage to the optic nerve, making timely diagnosis critical to prevent permanent vision loss. Although deep learning has demonstrated promising performance in automated glaucoma detection, existing approaches often overlook feature refinement, suffer from class imbalance, and rely on individual classifiers that limit prediction robustness. To address these challenges, this paper proposes a hybrid glaucoma detection framework that integrates attention-enhanced deep feature extraction with heterogeneous ensemble learning. Specifically, deep representations are extracted using InceptionV3 and subsequently refined by incorporating the Convolutional Block Attention Module (CBAM) to enhance discriminative retinal features. To improve classification robustness, the extracted features are classified using multiple machine learning models together with Single-Level Ensemble (SLE) and Double-Level Ensemble (DLE) strategies, while SMOTE combined with Tomek Links (SMOTE+TL) is employed to alleviate class imbalance. Furthermore, a systematic comparison of handcrafted, deep, and attention-enhanced deep feature representations is conducted. Experimental evaluation on two public retinal fundus datasets demonstrates that deep feature-based methods consistently outperform handcrafted feature-based methods, while the proposed attention-enhanced framework achieves the best overall performance. Furthermore, Grad-CAM visualizations confirm that the proposed model focuses on clinically relevant retinal regions, providing interpretable evidence on the model's prediction process.

CommentsAccepted and presented at 2026 IEEE International Conference on Biomedical Engineering, Computer and Information Technology for Health (BECITHCON)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑