arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

MIND:基于扩散Transformer的多模态意图驱动网络用于医学图像融合

MIND: Multimodal Intent-Driven Network via Diffusion Transformers for Medical Image Fusion

Yunzhan Fu, Xiangyu Shen, Yifei Sun, Yuhan Chen, Jian Wu, Hongxia Xu

arXiv 2607.28565首次发表:更新:

发表机构

Transvascular Implantation Devices Research Institute, Zhejiang University; Zhejiang University; Hangzhou Institute of Technology, Xidian University; Hangzhou Dianzi University(浙江大学血管内植入器械研究院; 浙江大学; 西安电子科技大学杭州研究院; 杭州电子科技大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对现有医学图像融合方法缺乏对诊断意图深度理解的问题,提出基于DiTs的MIND网络,通过BioMedGPT、多尺度潜在适配器和医学语义一致性损失优化,在多数据集上取得优异效果,可提升脑肿瘤分割精度并支持交互式融合。

AI 中文摘要

医学图像融合旨在整合不同成像模态的互补信息以辅助临床诊断。现有方法通常全局应用统一融合规则,缺乏对诊断意图和病理结构的深度理解。为解决这些局限,我们提出MIND,一种基于扩散Transformer(DiTs)的多模态意图驱动网络用于医学图像融合。具体而言,我们利用BioMedGPT从源图像生成意图驱动的融合文本,以病理感知的诊断意图指导融合过程。为应对DiTs中1D序列扁平化导致的2D空间连续性损失,我们设计了多尺度潜在适配器,该模块在序列化前显式提取源图像特征,通过严格的维度对齐将其注入网络以有效补充图像特征。为解决图像输出与诊断意图解耦导致的语义偏移,我们设计了医学语义一致性损失,该损失确保融合图像与融合文本之间的深度语义锁定,同时维持底层物理流形重建的稳定性。在Harvard、BraTS和GFP数据集上的综合实验表明,MIND实现了更优的融合质量,显著提升了下游脑肿瘤分割精度,且支持灵活的交互式融合,对意图驱动的智能临床决策支持系统具有重要应用前景。

英文摘要

Medical image fusion aims to integrate complementary information from diverse imaging modalities to support clinical diagnosis. Existing methods typically apply uniform fusion rules globally, lacking a deep understanding of diagnostic intents and pathological structures. To address these limitations, we propose MIND, a Multimodal Intent-Driven Network via Diffusion Transformers (DiTs) for medical image fusion. Specifically, we utilize BioMedGPT to generate intent-driven fusion texts from source images, guiding the fusion process with pathology-aware diagnostic intents. To combat the loss of 2D spatial continuity caused by 1D sequence flattening in DiTs, we design a Multi-scale Latent Adapter. This module explicitly extracts source image features before serialization, injecting them into the network via strict dimensional alignment to effectively supplement image features. To resolve the semantic shift caused by decoupling image outputs from diagnostic intents, we design a medical semantic consistency loss. This loss ensures deep semantic locking between fused images and fusion texts while maintaining the stability of the underlying physical manifold reconstruction. Comprehensive experiments on the Harvard, BraTS, and GFP datasets reveal that MIND delivers superior fusion quality, significantly improves downstream brain tumor segmentation accuracy, and enables flexible interactive fusion, holding significant promise for intent-driven intelligent clinical decision support systems.

Comments14pages, 14 figures, accepted by ACM MM2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑