arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2601.09879cs.CVcs.AI

MedVL-SAM2:一种统一的3D医学视觉-语言模型,用于多模态推理和基于提示的分割

MedVL-SAM2: A unified 3D medical vision-language model for multimodal reasoning and prompt-driven segmentation

  • Department of Biomedical Engineering, University of Florida(佛罗里达大学生物医学工程系)
  • Department of Radiology, University of Florida(佛罗里达大学放射学系)
  • Research Computing, University of Florida(佛罗里达大学研究计算中心)
  • Department of Medicine, University of Florida(佛罗里达大学医学系)
  • Department of Radiology, UC San Francisco(旧金山大学放射学系)

机构由 AI 辅助整理,请以论文原文为准。

Yang Xing, Jiong Wu, Savas Ozdemir, Ying Zhang, Yang Yang, Wei Shao, Kuang Gong

中文总结 AI 辅助

MedVL-SAM2是一种统一的3D医学多模态模型,通过联合训练实现报告生成、VQA和多任务分割的高性能表现。

中文摘要 AI 辅助

近年来,医学视觉-语言模型(VLMs)在图像层面以文本为中心的任务上取得了显著进展,如报告生成和视觉问答(VQA)。然而,在3D医学VLM中实现细粒度的视觉定位和体积空间推理仍具有挑战性,尤其是在试图在一个通用框架中统一这些能力时。为了解决这一挑战,我们提出了MedVL-SAM2,一种统一的3D医学多模态模型,同时支持报告生成、VQA和多范式分割,包括语义、指称和交互分割。MedVL-SAM2通过为3D医学影像量身定制的架构整合图像层面的推理和像素层面的感知,并结合基于SAM2的体积分割模块,以实现精确的多粒度空间推理。该模型通过多阶段管道进行训练:首先在大规模的3D CT图像-文本对语料库上预训练,以对齐体积视觉特征与放射学-语言嵌入。然后使用综合的3D CT分割数据集与语言理解和分割目标进行联合优化。这种联合训练使模型能够通过语言、点或框提示进行灵活交互,从而将高级视觉推理与空间精确定位统一起来。我们的统一架构在报告生成、VQA和多种3D分割任务上均实现了最先进的性能。广泛的分析进一步表明,该模型提供了可靠的3D视觉定位、可控的交互分割和稳健的跨模态推理,证明了在统一的3D医学VLM中可以同时实现高级语义推理和精确的3D定位。

英文摘要

Recent progress in medical vision-language models (VLMs) has achieved strong performance on image-level text-centric tasks such as report generation and visual question answering (VQA). However, achieving fine-grained visual grounding and volumetric spatial reasoning in 3D medical VLMs remains challenging, particularly when aiming to unify these capabilities within a single, generalizable framework. To address this challenge, we proposed MedVL-SAM2, a unified 3D medical multimodal model that concurrently supports report generation, VQA, and multi-paradigm segmentation, including semantic, referring, and interactive segmentation. MedVL-SAM2 integrates image-level reasoning and pixel-level perception through a cohesive architecture tailored for 3D medical imaging, and incorporates a SAM2-based volumetric segmentation module to enable precise multi-granular spatial reasoning. The model is trained in a multi-stage pipeline: it is first pre-trained on a large-scale corpus of 3D CT image-text pairs to align volumetric visual features with radiology-language embeddings. It is then jointly optimized with both language-understanding and segmentation objectives using a comprehensive 3D CT segmentation dataset. This joint training enables flexible interaction via language, point, or box prompts, thereby unifying high-level visual reasoning with spatially precise localization. Our unified architecture delivers state-of-the-art performance across report generation, VQA, and multiple 3D segmentation tasks. Extensive analyses further show that the model provides reliable 3D visual grounding, controllable interactive segmentation, and robust cross-modal reasoning, demonstrating that high-level semantic reasoning and precise 3D localization can be jointly achieved within a unified 3D medical VLM.

↑