CoMLP:用于医学图像细粒度跨模态信息融合的协作门控多层感知机
CoMLP: Cooperatively-Gated MLPs for Fine-Grained Cross-Modal Information Fusion in Medical Image Segmentation
浏览论文内容
中文总结 AI 辅助
本研究提出CoMLP协作门控多层感知机模块,通过协作交叉门控实现异源信息的细粒度跨模态融合,在5个医学分割基准上优于现有最优方法,为医学图像分割提供了MLP交互的有效替代方案。
中文摘要 AI 辅助
多模态医学图像与临床报告为医学图像分割提供互补的解剖、功能及语义信息,有效利用这些异源信息需保留细微空间细节且捕捉跨模态语义依赖的细粒度跨模态信息融合。现有融合方法常依赖交叉注意力,其计算负担随空间分辨率快速增长,难以在高分辨率特征图上实现密集跨模态交互,尤其对体积型医学图像而言。本研究提出CoMLP,一种用于医学图像细粒度跨模态信息融合的协作门控多层感知机模块。CoMLP通过协作交叉门控建模跨模态依赖,该门控构建于互补的区域与空洞多层感知机交互之上,以捕捉局部与全局跨模态依赖。我们进一步开发多源融合架构,其中CoMLP既执行成像模态间的图像间融合,又执行视觉特征与文本报告间的视觉-语言融合,使异源信息无需依赖密集交叉注意力即可整合。在涵盖2D/3D图像、临床报告、多种成像模态及不同解剖区域的5个医学分割基准上开展的大量实验表明,该方法相较于现有最优的多模态及语言引导分割方法取得一致提升;消融研究进一步显示,高空间分辨率下的细粒度交互及互补的局部-全局融合是性能提升的关键。这些结果表明,基于多层感知机的交互作为医学图像分割中细粒度跨模态信息融合的有效替代方案具有潜力。
英文摘要
Multi-modal medical images and clinical reports provide complementary anatomical, functional, and semantic information for medical image segmentation. Effectively exploiting these heterogeneous sources requires fine-grained cross-modal information fusion that preserves subtle spatial details while capturing semantic dependencies across modalities. Existing fusion approaches frequently rely on cross-attention, whose computational burden increases rapidly with spatial resolution, making dense cross-modal interaction difficult on high-resolution feature maps, particularly for volumetric medical images. In this work, we propose CoMLP, a cooperatively-gated MLP module for fine-grained cross-modal information fusion in medical image segmentation. CoMLP models cross-modal dependencies through cooperative cross-gating, built upon complementary regional and dilated MLP interactions, to capture local and global cross-modal dependencies. We further develop a multi-source fusion architecture in which CoMLP performs both inter-image fusion across imaging modalities and vision-language fusion between visual features and textual reports, enabling heterogeneous information to be integrated without relying on dense cross-attention. Extensive experiments on five medical segmentation benchmarks, covering 2D/3D images, clinical reports, multiple imaging modalities, and diverse anatomical regions, demonstrate consistent improvements over state-of-the-art multi-modal and language-guided segmentation methods. Ablation studies further show that fine-grained interaction at high spatial resolutions and complementary local-global fusion are critical to the performance gains. These results demonstrate the potential of MLP-based interaction as an effective alternative for fine-grained cross-modal information fusion in medical image segmentation.
发表机构
- Institute of Translational Medicine, Shanghai Jiao Tong University(上海交通大学转化医学研究院)
- Zhongguancun Academy(中关村学院)
- Zhongguancun Institute of Artificial Intelligence(中关村人工智能研究院)
- School of Computer Science, the University of Sydney(悉尼大学计算机学院)
- China United Network Communications Group Co., Ltd.(中国联合网络通信集团有限公司)
机构由 AI 辅助整理,请以论文原文为准。