arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2512.21372eess.IVcs.CV

一种基于图增强知识蒸馏的双流视觉Transformer与区域感知注意力的胃肠道疾病分类可解释AI方法

A Graph-Augmented knowledge Distillation based Dual-Stream Vision Transformer with Region-Aware Attention for Gastrointestinal Disease Classification with Explainable AI

  • Department of Computer Science and Engineering(计算机科学与工程系)

机构由 AI 辅助整理,请以论文原文为准。

Md Assaduzzaman, Nushrat Jahan Oyshi, Eram Mahamud

更新

AI总结:

本文提出一种结合图增强知识蒸馏的双流视觉Transformer与区域感知注意力机制,用于胃肠道疾病分类,通过软标签蒸馏实现高效且准确的诊断,同时保证模型的可解释性。

AI中文摘要:

胃肠道疾病从内窥镜和组织病理学影像的准确分类在医学诊断中仍是一个重大挑战,主要由于数据量庞大和类别间视觉差异细微。本研究提出了一种混合双流深度学习框架,基于教师-学生知识蒸馏,其中高容量教师模型集成了Swin Transformer的全局上下文推理与Vision Transformer的局部细粒度特征提取。学生网络采用紧凑的Tiny-ViT结构,通过软标签蒸馏继承教师的语义和形态知识,实现效率与诊断准确性的平衡。两个精心挑选的无线胶囊内镜数据集被用于确保类别平衡并防止样本间偏差。所提框架在数据集1和2上分别达到0.9978和0.9928的准确率,平均AUC为1.0000,表明几乎完美的判别能力。使用Grad-CAM、LIME和Score-CAM的可解释性分析证实,模型的预测基于临床显著的组织区域和病理相关形态学线索,验证了框架的透明性和可靠性。Tiny-ViT在计算复杂度较低的情况下,诊断性能与基于Transformer的教师模型相当,同时推理更快,适合资源受限的临床环境。总体而言,所提框架提供了一种稳健、可解释且可扩展的AI辅助胃肠道疾病诊断解决方案,为未来兼容临床实践的智能内窥镜筛查铺平道路。

英文摘要:

The accurate classification of gastrointestinal diseases from endoscopic and histopathological imagery remains a significant challenge in medical diagnostics, mainly due to the vast data volume and subtle variation in inter-class visuals. This study presents a hybrid dual-stream deep learning framework built on teacher-student knowledge distillation, where a high-capacity teacher model integrates the global contextual reasoning of a Swin Transformer with the local fine-grained feature extraction of a Vision Transformer. The student network was implemented as a compact Tiny-ViT structure that inherits the teacher's semantic and morphological knowledge via soft-label distillation, achieving a balance between efficiency and diagnostic accuracy. Two carefully curated Wireless Capsule Endoscopy datasets, encompassing major GI disease classes, were employed to ensure balanced representation and prevent inter-sample bias. The proposed framework achieved remarkable performance with accuracies of 0.9978 and 0.9928 on Dataset 1 and Dataset 2 respectively, and an average AUC of 1.0000, signifying near-perfect discriminative capability. Interpretability analyses using Grad-CAM, LIME, and Score-CAM confirmed that the model's predictions were grounded in clinically significant tissue regions and pathologically relevant morphological cues, validating the framework's transparency and reliability. The Tiny-ViT demonstrated diagnostic performance with reduced computational complexity comparable to its transformer-based teacher while delivering faster inference, making it suitable for resource-constrained clinical environments. Overall, the proposed framework provides a robust, interpretable, and scalable solution for AI-assisted GI disease diagnosis, paving the way toward future intelligent endoscopic screening that is compatible with clinical practicality.

↑