MagViT:用于乳腺组织病理学的可解释多倍率Transformer及患者级模型选择
MagViT: Interpretable Multi-Magnification Transformers with Patient-Level Model Selection for Breast Histopathology
- Dhaka International University(达卡国际大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文提出可解释多倍率Transformer框架MagViT,通过患者级五折交叉验证选择最优分支,在BreakHis、BUSI、IDC数据集上取得良好表现,且可解释性佳,提升乳腺组织病理学分类的稳健性。
AI中文摘要:
乳腺癌是全球女性最常见的癌症类型之一,快速检测与早期治疗可阻止其进展至更复杂阶段并抑制其向身体其他部位扩散。组织病理学图像分类是癌症检测中最常见的任务,因其在分析细胞数据方面具有稳健性。乳腺组织病理学分类需同时处理多尺度组织形态学,以及源域之外与临床相关的泛化问题。本文提出MagViT,这是一种可解释的多倍率Transformer框架,具备尺度门控融合与患者级模型选择功能。该模型使用BreakHis的四种倍率(40X、100X、200X、400X),通过ViT主干网络提取各尺度表征,并通过可学习门(可屏蔽缺失的尺度)将这些表征融合。研究采用固定种子的患者级五折交叉验证,并与三个架构分支进行比较;由于最精确的分支具备最强的患者级准确率且保留最简单的融合路径,因此被选为最终模型。在BreakHis数据集上,该架构实现的平均图像准确率为0.9191,平均患者准确率为0.9643,平均宏F1值为0.9042。外部迁移实验在BUSI(图像准确率0.8306,宏F1值0.7480,患者准确率0.8291)和IDC(图像准确率0.8577,宏F1值0.8191,患者准确率0.8372)的受控适配设置下,提供了跨数据集泛化的初步证据。Grad-CAM可视化表明,该模型在各倍率下均聚焦于具有诊断意义的重要区域。相较于此前以ViT为核心的BreakHis研究,本研究强调在可复现协议下的患者级选择与跨数据集稳健性。
英文摘要:
Breast cancer is one of the most common types of cancer among women around the world. Rapid detection and early treatment can hinder its progress to more complex stages and can impede its spread to other parts of the body. Histopathological image classification is the most common task in cancer detection due to its robustness in analyzing cellular data. Breast histopathology classification requires handling both multi-scale tissue morphology and clinically relevant generalization beyond the source domain. This paper presents MagViT, an interpretable multi-magnification transformer framework with scale-gated fusion and patient-level model selection. The model uses four BreakHis magnifications (40X, 100X, 200X, 400X) and extracts per-scale representations with a ViT backbone, and combines them via a learnable gate that masks missing scales. Patient-level five-fold cross-validation with a fixed seed has been run and compared with three architectural branches. The most accurate branch is then selected as the final model due to the strongest patient-level accuracy while retaining the simplest fusion pathway. On BreakHis, our architecture achieves a mean image accuracy of 0.9191, a mean patient accuracy of 0.9643, and a mean macro-F1 of 0.9042. External transfer experiments provide preliminary evidence of cross-dataset generalization under controlled adaptation settings on BUSI (image accuracy 0.8306, macro-F1 0.7480, patient accuracy 0.8291) and IDC (image accuracy 0.8577, macro-F1 0.8191, patient accuracy 0.8372). Grad-CAM visualization indicates that the model focuses on diagnostically significant and meaningful regions across magnifications. Relative to prior ViT-centered BreakHis work, this study emphasizes patient-level selection and cross-dataset robustness under a reproducible protocol.