MedVL-SAM2: A unified 3D medical vision-language model for multimodal reasoning and prompt-driven segmentation
MedVL-SAM2:一种统一的3D医学视觉-语言模型,用于多模态推理和基于提示的分割
机构 * Department of Biomedical Engineering, University of Florida(佛罗里达大学生物医学工程系) ; Department of Radiology, University of Florida(佛罗里达大学放射学系) ; Research Computing, University of Florida(佛罗里达大学研究计算中心) ; Department of Medicine, University of Florida(佛罗里达大学医学系) ; Department of Radiology, UC San Francisco(旧金山大学放射学系)
专题命中 视觉定位与Grounding :vision-language model(title,abstract);VLM(abstract);visual reasoning(abstract);visual question answering(abstract)
AI总结 MedVL-SAM2是一种统一的3D医学多模态模型,通过联合训练实现报告生成、VQA和多任务分割的高性能表现。