arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

学习看向何处:心脏MRI视觉语言模型的解剖学引导与注意力导向

Learning Where to Look: Anatomical Grounding and Guided Attention for Cardiac MRI Vision-Language Models

Bangwei Guo, Xiao Chen, Boris Mailhe, Jia Yao, Yiqing Wang, Ankush Mukherjee, Yikang Liu, Zheyuan Zhang, Hang Yu, Terrence Chen, Shanhui Sun

arXiv 2609.39899首次发表:更新:

发表机构

Rutgers University; United Imaging Intelligence; University of Texas Southwestern Medical Center; Duke University(罗格斯大学; 联影智能; 德克萨斯大学西南医学中心; 杜克大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对心脏MRI视觉语言模型缺乏解剖定位监督的问题,提出解剖学引导预训练与CARA注意力机制,构建大规模数据集,提升临床评估与区域定位能力。

AI 中文摘要

心脏磁共振成像(CMR)能够评估心脏解剖结构、心室功能和心肌组织特征。临床医生通过识别心脏结构并聚焦于与每个临床问题相关的区域来解读这些图像,这促使了解剖学引导的视觉语言模型(VLM)的发展。然而,针对解剖定位和临床问答的CMR特定监督仍然有限。为解决这一差距,我们通过解剖学引导和注意力导向研究细粒度的CMR视觉问答。我们构建了128,915个解剖学引导对和42,799个临床问答对,涵盖短轴电影、晚期钆增强和长轴电影。这些数据集支持解剖识别、定位和临床评估,无需为单个训练图像提供配对报告。为了帮助模型学习看向何处,我们引入了心脏解剖路由注意力(CARA),它根据问题选择预测的解剖先验,并以学习到的任务特定强度引导解码器注意力。将解剖学引导预训练与CARA结合,得到我们的模型CARA-VL。实验表明,CARA-VL在CMR成像设置中的临床评估和区域定位方面表现出优势,并具有向外部临床队列推广的良好泛化能力。总之,我们的数据和方法为研究和推进VLM中的心脏视觉理解提供了一个实用框架。我们将于发表后发布源自公共数据集的问答数据。

英文摘要

Cardiac magnetic resonance imaging (CMR) enables assessment of cardiac anatomy, ventricular function, and myocardial tissue characteristics. Clinicians interpret these images by identifying cardiac structures and focusing on the regions relevant to each clinical question, motivating anatomically guided vision-language models (VLMs). Yet CMR-specific supervision for anatomical localisation and clinical question answering remains limited. To address this gap, we investigate fine-grained CMR visual question answering through anatomical grounding and guided attention. We construct 128,915 anatomical-grounding and 42,799 clinical QA pairs across short-axis cine, late gadolinium enhancement, and long-axis cine. These datasets support anatomical recognition, localisation, and clinical assessment without requiring paired reports for individual training images. To help the model learn where to look, we introduce Cardiac Anatomy-Routed Attention (CARA), which selects predicted anatomical priors according to the question and guides decoder attention with learned task-specific strengths. Combining anatomical grounding pretraining with CARA yields our model, CARA-VL. Experiments demonstrate CARA-VL's strengths in clinical assessment and regional localisation across CMR imaging settings, with promising generalization to an external clinical cohort. Together, our data and method provide a practical framework for studying and advancing cardiac visual understanding in VLMs. We will release the QA data derived from public datasets upon publication.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑