arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.13309cs.CV

基础模型在纵向MRI疾病进展推理中的表现如何?

How Good are Foundation Models in Longitudinal MRI Disease Progression Reasoning?

Wafa Al Ghallabi, Ritesh Thawkar, Sara Ghaboura, Omkar Thawakar, Numan Saeed, Dana Al Nuaimi, Ajnas Alkatheeri, Salman Khan, Fahad Shahbaz Khan

首次发表
浏览论文内容

中文总结 AI 辅助

该研究推出Time-Aware Multi-View MRI基准,评估16个视觉-语言模型在纵向MRI疾病进展推理中的表现,发现模型在变化方向识别等方面存在缺陷,多视角输入对模型性能有特定影响。

中文摘要 AI 辅助

磁共振成像(MRI)解读是临床决策的基础,要求放射科医生整合不同时间点多视角的解剖平面,同时精准定位间隔变化。然而,现有的视觉-语言基准仍局限于单时间点、单视角的解读,无法捕捉放射实践中必需的时空推理能力。我们推出Time-Aware Multi-View MRI Benchmark(时间感知多视角MRI基准),这一评估框架整合了多视角解剖输入、纵向扫描的时间推理以及结构化定位引导。该基准包含3920个经专家验证的问答对,来自890名患者的3200多个纵向MRI时间点,涵盖7个临床队列,涉及胶质母细胞瘤、神经退行性疾病、前庭神经鞘瘤和脑转移瘤,采用开放式、多项选择和二元格式,要求模型识别变化最大的解剖区域、描述跨序列和视角的进展,并提供指定边界、成像特征和混杂因素的结构化引导。对16个视觉-语言模型的实验显示,这些模型具备一定的时间对齐能力,但在变化方向识别和体积量化方面存在系统性缺陷,而多视角输入可提升紧凑架构的空间定位能力,却会降低其时间推理能力。我们的基准为评估进展跟踪、间隔变化定位和时间排序提供了系统框架,这些对临床部署至关重要。代码、评估拆分和数据集可在该网址获取。

英文摘要

Magnetic Resonance Imaging (MRI) interpretation is fundamental to clinical decision-making, requiring radiologists to integrate multi-view anatomical planes across sequential timepoints while precisely localizing interval changes. However, existing vision-language benchmarks remain confined to single-timepoint, single-view interpretation, failing to capture the temporal-spatial reasoning essential to radiologic practice. We introduce the Time-Aware Multi-View MRI Benchmark, an evaluation framework unifying multi-view anatomical input, temporal reasoning across longitudinal scans, and structured localization guidance. The benchmark comprises 3,920 expert-verified question-answer pairs derived from 890 patients across over 3,200 longitudinal MRI timepoints, drawn from seven clinical cohorts covering glioblastoma, neurodegeneration, vestibular schwannoma, and brain metastases, in open-ended, multiple-choice, and binary formats, requiring models to identify anatomical regions of maximal change, characterize progression across sequences and views, and provide structured guidance specifying boundaries, imaging features, and confounders. Experiments across 16 vision-language models reveal moderate temporal alignment but systematic failure on change direction recognition and volumetric quantification, while multi-view inputs improve spatial localization yet degrade temporal reasoning in compact architectures. Our benchmark provides a systematic framework for evaluating progression tracking, interval change localization, and temporal ordering, which are essential for clinical deployment. Code, evaluation splits, and the dataset are available at: https://github.com/wafaAlghallabi/Time-Aware-MRI.

发表机构

  • Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学)
  • Department of Health Abu Dhabi(阿布扎比卫生部)
  • Fatima College of Health Sciences(法蒂玛健康科学学院)
  • Linköping University(林雪平大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑