arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

CineMR:用于定量心脏MRI评估的工具集成视觉语言推理

CineMR: Tool-Integrated Vision-Language Reasoning for Quantitative Cardiac MRI Assessment

Kunyang Li, Hai Nguyen, Joshua Lowe, Chenguang Zhao, Peace C. Madueme, Mehdi Hedjazi Moghari, Mubarak Shah, Pegah Khosravi, Yuzhang Shang

arXiv 2610.01166首次发表:更新:

发表机构

Institute for Artificial Intelligence, University of Central Florida; Nemours Cardiac Center, Nemours Children’s Hospital, Florida; Children’s Heart Center, WVU Golisano Children’s, West Virginia; Department of Clinical Sciences, College of Medicine, University of Central Florida(中佛罗里达大学人工智能研究所; 佛罗里达州内穆尔斯儿童医院内穆尔斯心脏中心; 西弗吉尼亚大学戈利萨诺儿童医院儿童心脏中心; 中佛罗里达大学医学院临床科学系)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

CineMR是一种工具增强的视觉语言模型,通过调用心脏图像分析工具并整合输出,实现定量CMR评估,在基准上显著提升性能,强调了可靠工具使用的重要性。

AI 中文摘要

心血管磁共振(CMR),包括电影成像,是无创评估心脏形态和心室功能的参考标准。电影CMR解读将定性视觉评估与心室容积、射血分数、心肌质量、室壁厚度和局部室壁运动的定量测量相结合。当前的医学视觉语言模型(VLM)无法在没有分析工具的情况下从多维电影图像中可靠地得出定量测量结果。我们提出了CineMR,一种工具增强的VLM,它调用心脏图像分析工具并将其输出整合到交错推理中,用于定量CMR评估。我们还构建了一个多队列视觉问答基准,涵盖定量指标提取、多类诊断和鉴别诊断,以及用于分割、相位选择、容积测量、形态测量和局部室壁运动分析的工具。CineMR通过监督微调(SFT)在工具交互轨迹上进行训练,随后使用带有条件工具使用奖励的组相对策略优化(GRPO)。在多队列电影CMR基准上,CineMR实现了35.9%的pass@1和58.9%的pass@4,而Qwen3-VL-8B骨干网络的pass@1为1.5%,LLaVA-Med v1.5和MedGemma-4B的pass@1分别为0.0%和7.0%。GRPO后正确的工具调用达到99.8%,高于SFT后的78.9%。实时工具输出将心室测量准确性提高了20.4%至23.7%,而移除所有工具将pass@1从35.9%降至27.9%。这些结果强调了可靠工具使用对定量电影CMR推理的重要性,并支持CineMR作为辅助心脏图像评估的一种有前景的方法。代码、基准资源和模型权重可在以下https URL获取。

英文摘要

Cardiovascular magnetic resonance (CMR), including cine imaging, is a reference standard for the noninvasive assessment of cardiac morphology and ventricular function. Cine CMR interpretation integrates qualitative visual assessment with quantitative measurements of ventricular volumes, ejection fraction, myocardial mass, wall thickness, and regional wall motion. Current medical vision-language models (VLMs) cannot reliably derive quantitative measurements from multidimensional cine images without analysis tools. We present CineMR, a tool-augmented VLM that invokes cardiac image-analysis tools and integrates their outputs into interleaved reasoning for quantitative CMR assessment. We also construct a multi-cohort visual question answering benchmark covering quantitative metric extraction, multiclass diagnosis, and differential diagnosis, together with tools for segmentation, phase selection, volumetry, morphometry, and regional wall motion analysis. CineMR is trained with supervised fine-tuning (SFT) on tool-interaction traces followed by Group Relative Policy Optimization (GRPO) with conditional tool-use rewards. On the multi-cohort cine CMR benchmark, CineMR achieves 35.9% pass@1 and 58.9% pass@4, compared with 1.5% pass@1 for the Qwen3-VL-8B backbone and 0.0% and 7.0% pass@1 for LLaVA-Med v1.5 and MedGemma-4B, respectively. Correct tool invocation reaches 99.8% after GRPO, up from 78.9% after SFT. Live tool outputs improve ventricular measurement accuracy by 20.4--23.7% over direct model predictions, and removing all tools reduces pass@1 from 35.9% to 27.9%. These results highlight the importance of reliable tool use for quantitative cine CMR reasoning and support CineMR as a promising approach for assistive cardiac image assessment. Code, benchmark resources, and model weights are available at https://github.com/AI-MIND-Lab/CineMR.

CommentsCode, benchmark resources, and model weights are available at https://github.com/AI-MIND-Lab/CineMR

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑