AI 中文总结
本文针对现有3D视觉语言模型无法满足多参数3D MRI跨模态协同推理需求的问题,提出Mr3D-VL模型,通过特定设计实现性能提升,在多项任务上优于同规模的领域及通用模型。
AI 中文摘要
多参数磁共振成像(mpMRI)是脑肿瘤诊断与治疗的核心手段,但当前AI模型存在关键局限:缺乏自然语言交互能力与可解释性,阻碍了临床所需的空间信息整合与跨模态推理。核心挑战源于不同模态间物理意义差异显著、扫描间隔导致的空间错位,以及胶质瘤分级等任务中对复杂多特征解释的需求。尽管视觉语言模型(VLMs)在跨模态理解方面展现潜力,但现有方法主要聚焦2D图像建模,忽视对3D体积空间的直接感知;虽已有3D VLMs被提出用于3D CT成像的报告生成与特征对齐,但mpMRI应用需跨多个成像模态的协同推理,这一需求尚未被现有方案满足。为解决该问题,本文提出Mr3D-VL,一种专为多参数3D MRI设计的视觉语言基础模型,其拥有40亿参数,采用无监督预训练的共享3D编码器与4D旋转位置嵌入实现双模态-空间整合;跨模态投影层采用多分辨率特征植入策略以增强不同分辨率下的特征感知。实验结果显示,该模型在文本生成任务上显著优于现有4B/7B/30B规模的领域特定及通用模型,报告生成任务的BERTScore达0.856,问答准确率为0.713,多选题准确率为0.912。
英文摘要
Multi-parametric magnetic resonance imaging (mpMRI) is a cornerstone for brain tumor diagnosis and treatment, yet current AI models face critical limitations: their lack of natural language interaction and interpretability impedes spatial information integration and cross-modal reasoning required clinically. Key challenges arise from significant physical meaning differences across modalities, spatial misalignment due to scan intervals, and the need for complex multi-feature interpretation in tasks like glioma grading. While visual-language models (VLMs) show promise in cross-modal understanding, existing methods focus mainly on 2D image modeling, neglecting direct perception of 3D volumetric space. Although 3D VLMs have been proposed for report generation and feature alignment in 3D CT imaging, mpMRI applications demand collaborative inference across multiple imaging modalities-a requirement unmet by current solutions. To address this, we introduce Mr3D-VL, a dedicated visual-language foundation model for multi-parametric 3D MRI. With 4 billion parameters, it employs an unsupervised pre-trained shared 3D encoder and 4D rotational positional embedding for dual modality-spatial integration. Its cross-modal projection layer uses a multi-resolution feature implantation strategy to enhance feature perception across resolutions. Experimental results show significant improvements over existing 4B/7B/30B domain-specific and general-purpose models in text generation tasks, achieving a BERTScore of 0.856 for report generation, with question-answering accuracy at 0.713 and multiple-choice accuracy at 0.912.