发表机构
University of Central Florida; University of Rochester; Harvard Medical School; University of Virginia(中佛罗里达大学; 罗切斯特大学; 哈佛医学院; 弗吉尼亚大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
UniMod框架通过要求各模态独立预测诊断缓解多模态医学诊断的捷径学习,添加跨模态与模态内对齐,在两个数据集上优于基线方法,还可扩展至多标签5分类诊断。
AI 中文摘要
结合医学图像与临床文本的多模态学习在疾病诊断中具有应用前景,但标准多模态训练会导致捷径学习:模型会利用更容易学习的模态(如文本中的诊断线索),而忽略更难学习的特征(如细微的视觉模式)。本文提出UniMod框架,通过要求每个模态独立预测诊断来缓解捷径学习,同时监督仅图像、仅文本及多模态分类,迫使各模态提取诊断特征;还添加跨模态对齐以实现知识迁移,以及针对相同诊断患者的模态内监督对比对齐。在Harvard-Glaucoma数据集上,UniMod的AUC达0.850,较OGM-GE和Gradient Blending高出1.6-1.8%;在CheXpert Plus数据集上,其AUC达0.966,超出上述两种方法5%以上;UniMod无需修改架构即可扩展至多标签5分类诊断,较CGGM的平均AUC提升0.097。
英文摘要
Multi-modal learning combining medical images and clinical text is promising for disease diagnosis. However, standard multi-modal training leads to shortcut learning: models exploit the easier modality (e.g., diagnostic cues in text) while neglecting harder-to-learn features (e.g., subtle visual patterns). We propose UniMod, a framework that mitigates shortcut learning by requiring each modality to predict the diagnosis on its own. It supervises image-only, text-only, and multi-modal classification simultaneously, so each modality must extract diagnostic features. We add cross-modality alignment for knowledge transfer and within-modality supervised contrastive alignment over same-diagnosis patients. On Harvard-Glaucoma, UniMod reaches 0.850 AUC, outperforming OGM-GE and Gradient Blending by 1.6-1.8%; on CheXpert Plus, it reaches 0.966 AUC, surpassing them by over 5%. UniMod also extends to 5-class multi-label diagnosis without architectural change, improving mean AUC by 0.097 over CGGM.
CommentsAccepted to ACM Multimedia 2026 (MM '26). 10 pages, 7 figures, 5 tables. Code: https://github.com/futurespyhi/UniMod