发表机构
University of Toronto; The Hospital for Sick Children; Vector Institute; University Hospital Augsburg(多伦多大学; 病童医院; 向量研究所; 奥格斯堡大学医院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对3D医学成像中假阴性问题,提出多模态语义感知对比学习框架MseaCL,通过纳入放射学报告语义相似性作指导信号,经实验验证该框架用于预训练可提升下游任务表现,如儿科脑肿瘤分子分类AUC至少增22.6%。
AI 中文摘要
多模态对比学习(CL)在跨数据模态对齐表示和改善下游任务方面表现出色,尤其在医疗保健领域。它通过最小化匹配(正)数据模态之间的距离,同时最大化不匹配(负)样本之间的距离来工作。传统CL框架通常假设数据批次内基于实例的对应关系,将所有非配对样本视为负样本。然而,这种假设在医学环境中常常不成立,样本可能共享高级语义属性,导致假阴性,降低表示质量。本文提出了多模态语义感知对比学习(MseaCL),这是一种在3D脑磁共振成像(MRI)扫描和放射学报告的儿科队列上训练的CL框架。该框架的目标是通过纳入放射学报告之间的语义相似性,作为学习过程中的指导信号,减轻语义相似的假阴性样本的影响。我们的结果表明,将此框架用作预训练阶段可以在下游任务中取得显著改进,例如,儿科脑肿瘤分子分类的受试者操作特征曲线(AUC)下面积至少增加22.6%,证明了其在临床应用中实现更强大和语义对齐的多模态表示的潜力。
英文摘要
Multimodal Contrastive Learning (CL) has shown significant performance in aligning representations across various data modalities and improving downstream tasks, especially in healthcare. It works by minimizing the distance between matched (positive) data modalities, while maximizing the distance between mismatched (negative) samples. Traditional CL frameworks typically assume instance-based correspondence within data batches, treating all non-paired samples as negatives. However, this assumption often fails in medical settings, where samples may share high-level semantic attributes, leading to false negatives that degrade representation quality. In this paper, we propose Multimodal Semantic-Aware Contrastive Learning (MseaCL), a CL framework trained on a pediatric cohort of 3D brain magnetic resonance imaging (MRI) scans and radiology reports. The goal of this framework is to mitigate the impact of semantically similar false negative samples by incorporating semantic similarity between radiology reports, as a guiding signal during the learning process. Our results indicate that applying this framework as a pretraining stage can achieve notable improvements in downstream tasks, e.g., at least a 22.6\% increase in the area under the receiver operating characteristic curve (AUC) of pediatric brain tumor molecular classification, demonstrating its potential for more robust and semantically aligned multimodal representations in clinical applications.