arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.26856cs.CVcs.AI

从推理到像素:用于VQA和分割的接地医学多模态大语言模型

From Reasoning to Pixels: Grounded Medical Multimodal LLMs for VQA and Segmentation

  • Nanjing University of Science and Technology(南京理工大学)
  • Sungkyunkwan University(成均馆大学)

机构由 AI 辅助整理,请以论文原文为准。

Haowen Gu, Gensheng Pei, Junzhu Mao, Qiong Wang, Mingwu Ren, Yazhou Yao

AI总结:

针对现有医学多模态大语言模型缺乏像素级接地的问题,提出MedREAL框架,引入SARP与R2V机制,构建MedRAVS-13K数据集,在Med-VQA和分割任务上性能优于现有方法,为医学图像分析提供可解释框架。

AI中文摘要:

尽管多模态大语言模型(MLLMs)在医学视觉问答(Med-VQA)中已展现出令人印象深刻的性能,但它们对全局图像特征的依赖往往缺乏精确的像素级接地,从而限制了临床可信度。为弥合高级临床推理与空间定位之间的语义差距,我们提出了MedREAL(Medical Reasoning-driven Answering and Localization,医学推理驱动的问答与定位),这是一个将语言推理与空间接地无缝对齐的统一框架。具体而言,MedREAL引入了分割锚定推理池化(SARP),以直接从MLLM隐藏状态中的[SEG]标记中提取与任务相关的语义证据。此外,我们还提出了推理到视觉(R2V)融合机制,以将这些感知推理特征有效注入分割流水线,从而实现精确的掩码解码。为推动这一范式,我们构建了MedRAVS-13K,这是一个包含13824个经专家验证样本的综合数据集,涵盖四种不同的成像模态。大量实验表明,MedREAL显著优于现有最优方法,在基准评估中达到68.49%的gIoU和70.47%的cIoU。通过生成与文本诊断严格一致的证据掩码,MedREAL为推理驱动的医学图像分析提供了一个稳健、可解释的框架。

英文摘要:

Although Multimodal Large Language Models (MLLMs) have demonstrated impressive performance in Medical Visual Question Answering (Med-VQA), their reliance on global image features often lacks precise pixel-level grounding, thereby limiting clinical trustworthiness. To bridge the semantic gap between high-level clinical reasoning and spatial localization, we propose \textsc{\textsc{MedREAL}} (\textbf{Med}ical \textbf{RE}asoning-driven \textbf{A}nswering and \textbf{L}ocalization), a unified framework that seamlessly aligns linguistic reasoning with spatial grounding. Specifically, \textsc{MedREAL} introduces \textbf{S}eg \textbf{A}nchored \textbf{R}easoning \textbf{P}ooling (SARP) to distill task-relevant semantic evidence directly from \texttt{[SEG]} tokens within the MLLM's hidden states. Furthermore, a \textbf{R}easoning-to-\textbf{V}isual (R2V) fusion mechanism is proposed to effectively inject these reasoning-aware features into a segmentation pipeline for accurate mask decoding. To facilitate this paradigm, we construct MedRAVS-13K, a comprehensive dataset comprising 13,824 expertly validated samples across four diverse imaging modalities. Extensive experiments demonstrate that \textsc{MedREAL} significantly outperforms state-of-the-arts, achieving 68.49\% gIoU and 70.47\% cIoU on benchmark evaluations. By generating evidence masks that are strictly consistent with textual diagnoses, \textsc{MedREAL} provides a robust, interpretable framework for reasoning-driven medical image analysis.

补充信息

↑