From Reasoning to Pixels: Grounded Medical Multimodal LLMs for VQA and Segmentation
从推理到像素:用于VQA和分割的接地医学多模态大语言模型
机构 * Nanjing University of Science and Technology(南京理工大学) ; Sungkyunkwan University(成均馆大学)
专题命中 视觉问答 :MLLM(abstract,abstract_cn);visual question answering(abstract);grounding(abstract);multimodal large language model(abstract)
AI总结 针对现有医学多模态大语言模型缺乏像素级接地的问题,提出MedREAL框架,引入SARP与R2V机制,构建MedRAVS-13K数据集,在Med-VQA和分割任务上性能优于现有方法,为医学图像分析提供可解释框架。
Comments accepted by ECCV 2026