BrainTRACE:在脑部MRI临床推理中追踪纵向、多模态和体积证据
BrainTRACE: Tracing Longitudinal, Multimodal, and Volumetric Evidence in Brain MRI Clinical Reasoning
浏览论文内容
中文总结 AI 辅助
BrainTRACE是一个基于报告的基准,用于评估视觉语言模型在纵向脑部MRI解读中追踪多模态、体积证据的能力,包含7,273个实例,发现当前模型难以整合证据链。
中文摘要 AI 辅助
脑部MRI解读是一个纵向的临床推理问题:放射科医生比较系列研究,整合跨MRI序列的信息,在体积解剖结构中定位发现,并将这些证据转化为基于报告的评估。现有的医学VQA和3D成像基准涵盖了这一工作流程的重要部分,但通常通过孤立的图像、静态体积或未基于报告的答案风格来评估脑部MRI,从而掩盖了支持临床有效性的证据链中的失败。我们引入了BrainTRACE,一个基于报告的基准,用于评估视觉语言模型能否追踪纵向脑部MRI解读所需的证据结构。BrainTRACE包含从1,778名纵向患者、7,299项MRI研究和约29k个共配准的3D MRI序列体积中衍生的7,273个评分VQA实例。该基准按五个临床推理级别组织,从采集识别到病例级综合,并按证据需求涵盖纵向比较、基于报告的引用、多序列整合和体积空间证据。BrainTRACE支持与标准VLM接口兼容的渲染输入、3D证据条件,以及一个分解的病例推理轨道,该轨道审计纵向证据链中的六个步骤。对20种VLM配置的评估表明,当前系统能识别孤立的视觉线索,但很少能将它们组合成基于证据的纵向解释。我们发布了基准规范、评估列表、评分实现、评分细则和审计记录格式,以支持脑部MRI VLM评估的可复现进展。
英文摘要
Brain MRI interpretation is a longitudinal clinical reasoning problem: radiologists compare serial studies, integrate information across MRI sequences, localize findings within volumetric anatomy, and translate this evidence into report-grounded assessments. Existing medical VQA and 3D imaging benchmarks capture important parts of this workflow, but often evaluate brain MRI through isolated images, static volumes, or ungrounded report-style answers, thereby obscuring failures in the evidence chain that support clinical validity. We introduce BrainTRACE, a report-grounded benchmark for evaluating whether vision-language models can trace the evidence structure required for longitudinal brain MRI interpretation. BrainTRACE contains 7,273 scored VQA instances derived from 1,778 longitudinal patients, 7,299 MRI studies, and approximately 29k co-registered 3D MRI sequence volumes. The benchmark is organized by five levels of clinical reasoning, from acquisition recognition to case-level synthesis, and by evidence demands covering longitudinal comparison, report-grounded references, multi-sequence integration, and volumetric spatial evidence. BrainTRACE supports rendered inputs compatible with standard VLM interfaces, a 3D-evidence condition, and a decomposed case-reasoning track that audits six steps in a longitudinal evidence chain. Evaluation of 20 VLM configurations shows that current systems can identify isolated visual cues but rarely compose them into grounded longitudinal interpretations. We release the benchmark specification, evaluation lists, scoring implementation, scoring rubrics, and audit-record format to support reproducible progress in brain MRI VLM evaluation.
发表机构
- University of Texas Health Science Center at Houston(德克萨斯大学休斯顿健康科学中心)
- University of Alabama at Birmingham(阿拉巴马大学伯明翰分校)
- University of Pittsburgh(匹兹堡大学)
- The University of Texas MD Anderson Cancer Center(德克萨斯大学MD安德森癌症中心)
- University of Dublin, Trinity College(都柏林大学圣三一学院)
- Yale University(耶鲁大学)
机构由 AI 辅助整理,请以论文原文为准。