AI 中文总结
该研究开发了SPARC-Rad多模态基准数据集及评估流程,用于评估放射学VLMs的空间与解剖推理能力,为相关模型的开发与部署前评估提供支持。
AI 中文摘要
视觉语言模型(VLMs)在医学成像领域的评估愈发受到关注,但现有诸多基准侧重疾病分类、报告生成或通用视觉问答,而非放射学所需的空间与解剖推理能力。我们开发了临床放射学空间感知与解剖推理(SPARC-Rad)基准,这是一个人工整理的多模态基准数据集及评估流程,用于评估放射学VLMs的相关能力。SPARC-Rad包含300组图像-问题对,源自癌症影像档案(TCIA)的健康对照成像研究,涵盖腹部、胸部、乳腺、神经及肌肉骨骼类别的CT、MRI与放射影像。放射学受训人员手动设计并标注问题,以评估解剖结构识别、定位、侧别判断、区域识别、设备识别及结构间空间关系。该评估流程支持标准化提示、结构化输出收集、响应归一化、大语言模型(LLM)作为评判者打分、人工质量审核、二元正确性评分,以及按模态、解剖部位和推理类型开展的亚组分析。SPARC-Rad提供了一个可复用框架,用于评估VLMs能否以空间系统的形式为放射解剖学提供推理支持,助力未来模型开发、失败模式分析及部署前评估。
英文摘要
Vision-language models (VLMs) are increasingly being evaluated for medical imaging, but many available benchmarks emphasize disease classification, report generation, or broad visual question answering rather than the spatial and anatomical reasoning required for radiology. We developed the Spatial Perception and Anatomical Reasoning in Clinical Radiology (SPARC-Rad) Benchmark, a manually curated multimodal benchmark dataset and evaluation pipeline for assessing these capabilities in radiology VLMs. SPARC-Rad includes 300 image-question pairs derived from healthy control imaging studies in The Cancer Imaging Archive (TCIA), spanning CT, MRI, and radiography across the abdomen, chest, breast, neuro, and musculoskeletal categories. Radiology trainees manually designed and annotated questions to evaluate anatomical identification, localization, laterality, regional recognition, device identification, and inter-structure spatial relationships. The evaluation pipeline supports standardized prompting, structured output collection, response normalization, LLM-as-judge grading, human quality review, binary correctness scoring, and subgroup analysis by modality, anatomy, and reasoning type. SPARC-Rad provides a reusable framework for evaluating whether VLMs can provide reasoning for radiologic anatomy as a spatial system, supporting future model development, failure-mode analysis, and pre-deployment assessment.