arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.03991cs.CVcs.LG

掩蔽特权信息蒸馏用于缺失临床元数据下的多模态皮肤病变分类

Masked Privileged-Information Distillation for Multimodal Skin Lesion Classification Under Missing Clinical Metadata

Anirban Barua, Md Mahir Abrar Khan, Ayman Iktidar, Md. Sajjatul Islam

首次发表
浏览论文内容

中文总结 AI 辅助

提出掩蔽特权信息蒸馏框架,用完整元数据训练的教师监督掩蔽训练的小学生模型,在PAD-UFES-20上提升平衡准确率并实现优雅降级,适用于资源受限的即时护理场景。

中文摘要 AI 辅助

多模态皮肤病变分类结合临床图像与患者元数据以提高诊断准确性。然而,训练时可获得的完整元数据在部署时可能仅部分可用,且资源受限的环境还要求计算效率。我们通过一个特权信息蒸馏框架应对这些挑战,在该框架中,一个在完整元数据上训练的多模态教师模型监督一个在随机掩蔽临床字段上训练的9.2倍更小的学生模型。临床字段以从U(0,1)中抽取的每样本比率作为整体组进行掩蔽,因此一次训练运行覆盖完整的元数据可用性范围。融合是残差式的,元数据作为门控校正添加到无条件图像基础之上。在PAD-UFES-20数据集上,掩蔽训练下的蒸馏在每个可用性水平上相比交叉熵训练提高了平衡准确率,平均提高+4.7个百分点,而无掩蔽时仅提高+1.8个百分点。当元数据从完整减少到缺失时,掩蔽学生仅损失7.8个平衡准确率点,而同一学生在完整元数据上训练则损失36.6个点,这凸显了掩蔽训练在超越单纯蒸馏的优雅降级中的作用。Grad-CAM可视化进一步显示,随着元数据的撤除,掩蔽学生的注意力通常仍以病变为中心。所得到的紧凑模型面向临床元数据往往不完整的即时护理场景。

英文摘要

Multimodal skin lesion classification combines clinical images with patient metadata to improve diagnostic accuracy. However, complete metadata available during training may be only partially accessible at deployment, and resource-constrained settings additionally require computational efficiency. We address these challenges with a privileged-information distillation framework in which a multimodal teacher trained on complete metadata supervises a 9.2x smaller student trained with randomly masked clinical fields. Clinical fields are masked as whole groups at a per-sample rate drawn from U(0,1), so one training run covers the full metadata availability range. Fusion is residual, with metadata added as a gated correction to an unconditional image base. On the PAD-UFES-20 dataset, distillation under masked training improves balanced accuracy over cross-entropy training at every availability level, by an average of +4.7 points versus +1.8 points without masking. The masked student loses only 7.8 balanced-accuracy points as metadata decreases from complete to absent, compared with 36.6 points for the same student trained on complete metadata, highlighting the role of masked training in graceful degradation beyond distillation alone. Grad-CAM visualizations further show that the masked student's attention generally remains lesion-centered as metadata is withdrawn. The resulting compact model targets point-of-care settings, where clinical metadata is often incomplete.

发表机构

  • Chittagong Independent University (CIU)(吉大港独立大学)

机构由 AI 辅助整理,请以论文原文为准。

↑