发表机构
Indian Institute of Technology, Jodhpur(印度理工学院焦特布尔分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出多模态知识蒸馏框架,利用低秩融合训练教师模型,蒸馏至学生模型实现仅图像推理,在胃腺癌WSI分类中比现有方法平均准确率提升至少3.35%。
AI 中文摘要
胃腺癌(GA)是全球癌症相关死亡的主要原因之一,从全切片图像(WSIs)中进行准确的组织病理亚型分类对于有效的治疗规划至关重要。虽然整合病理报告文本与WSIs的多模态方法可以改善分类性能,但现有方法通常依赖于计算成本高昂的Transformer架构和大语言模型。我们提出了一种多模态知识蒸馏(MKD)框架,该框架结合了预训练的WSI图像编码器和临床文本编码器,利用低秩多模态融合(LMF)在训练期间高效地建模跨模态交互。每个WSI被表示为一个补丁包,并配有一个切片级别的诊断描述。教师模型学习用于亚型分类的融合图像-文本表示,而学生模型则蒸馏这些知识以实现仅基于图像的准确推理。我们在PatchGastric基准数据集上评估了我们的方法,在不依赖基于Transformer的融合、多任务学习或大语言模型的情况下,实现了比最先进方法至少高出3.35%的平均准确率。源代码可在该https URL获取。
英文摘要
Gastric adenocarcinoma (GA) is a leading cause of cancer-related mortality worldwide, and accurate histopathological subtype classification from whole-slide images (WSIs) is essential for effective treatment planning. While multimodal approaches that integrate pathology report text with WSIs can improve classification, existing methods often depend on computationally expensive transformer architectures and large language models. We propose a multimodal knowledge distillation (MKD) framework that combines a pretrained WSI image encoder and a clinical text encoder using Low-Rank Multimodal Fusion (LMF) to efficiently model cross-modal interactions during training. Each WSI is represented as a bag of patches paired with a slide-level diagnostic caption. The teacher model learns fused image-text representations for subtype classification, while the student model distills this knowledge to enable accurate image-only inference. We evaluate our method on the PatchGastric benchmark dataset and achieve at least 3.35% higher mean accuracy than state-of-the-art approaches, without relying on transformer-based fusion, multi-task learning, or large language models. The source code is available at https://github.com/helomelo1/MKD-LMF.