arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

MAGE:用于胃肿瘤分类的颜色不变和空间知识蒸馏

MAGE: Color-Invariant and Spatial Knowledge Distillation for Gastric Neoplasm Classification

Jiho Jun, Jeongwon Woo, Jaemin Song, Thanh Bong Nguyen, Dong-heon Yeon, Donghoon Kang, Jae-Myung Park, Sung-Jea Ko, Kwang-Hyun Uhm

arXiv 2607.12663首次发表:更新:

发表机构

Korea University; MEDAI; Seoul National University; Vietnam National University, Hanoi; The Catholic University of Korea, Seoul ST. Mary’s Hospital; Gachon University(韩国大学; MEDAI; 首尔国立大学; 越南河内国家大学; 韩国天主教大学首尔圣母医院; 加图立大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对胃肿瘤分类难题,提出MAGE框架,训练时引入辅助局部专家分支并采用双目标蒸馏策略,推理时可实时部署,实验证明该方法在性能和可解释性上优于现有方法。

AI 中文摘要

在内窥镜检查中准确区分胃腺瘤和癌对临床决策至关重要。但由于两类之间的高类间相似性和模糊边界,该任务极具挑战性。现有基于ROI的分类方法常受检测/分割误差传播和周围全局上下文丢失影响,全图像分类缺乏必要空间焦点,且神经网络易受特定领域纹理偏差影响。为此提出MAGE框架,训练时引入在肿瘤消色差视图上训练的辅助局部专家分支,采用双目标蒸馏策略,推理时可在无注释掩码的图像上运行。实验表明该方法显著优于现有方法,提供了更好的分类性能和可解释的注意力图。

英文摘要

Accurate differentiation between gastric adenoma and carcinoma during endoscopy is critical for clinical decision-making. Yet, this task is highly challenging due to high inter-class similarity and ambiguous boundaries between the two classes. Existing ROI-based classification methods often suffer from detection/segmentation error propagation and loss of surrounding global context. In contrast, full-image classification lacks the necessary spatial focus. Furthermore, we observe that deep neural networks gravitate towards domain-specific texture biases(e.g. bleeding, lighting artifacts), often causing models to predict based on spurious correlations instead of intrinsic morphological features. To address these limitations, we propose a novel framework, Masked Achromatic Guidance Expert (MAGE). During training, we introduce an auxiliary local expert branch trained on masked achromatic views of the neoplasm. By suppressing background context and color, this branch is forced to learn highly discriminative, purely structural features. We then employ a dual-objective distillation strategy, transferring both classification logits and spatial attention maps to provide implicit spatial supervision to the main branch that receives full WLI as input. This dual-objective distillation forces the model to ground its predictions in morphology rather than relying on shortcuts, while still retaining clinically relevant color cues. At inference time, our deployable model operates on images without annotated masks, ensuring real-time deployability . Extensive experiments on a clinical gastric endoscopy dataset show that our method significantly outperforms existing detection-based methodologies (e.g. YOLO) and classification-based methodologies (e.g. Swin-Transformer), providing not only superior classification performance but also interpretable attention maps for clinical reliability.

CommentsAccepted to MICCAI 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑