发表机构
Shenzhen Institute of Advanced Technology, Chinese Academy of Sciences; Pengcheng Laboratory; University of Chinese Academy of Sciences; Peking Union Medical College Hospital, Chinese Academy of Medical Sciences and Peking Union Medical College(中国科学院深圳先进技术研究院; 鹏城实验室; 中国科学院大学; 北京协和医学院附属北京协和医院(中国医学科学院北京协和医学院))
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究建立含3021个CCTA序列的四中心基准,开发临床结构化指标CSM$_{\text{CCTA}}$,发现CCTA训练的C2RG模型性能最优但仍有差距,通用模型输出多为无关报告,为CCTA报告生成评估提供标准化方案
AI 中文摘要
可靠评估自动化冠状动脉计算机断层血管造影(CCTA)报告生成,需要标准化的多中心基准及临床结构化指标。我们建立了一个包含3021个CCTA序列的四中心基准,这些序列来自818份患者-报告对,用于评估7种开源三维视觉-语言模型。我们开发了CSM$_{\text{CCTA}}$,一种用于CCTA报告评估的临床结构化指标,其患者、血管和节段级变量均根据临床指南定义。报告对在最精细的共享解剖学层面进行比较,不同临床成分的贡献根据专家评估进行加权。我们使用70个经专家评分的案例估计这些权重,并在非重叠的30个案例集中评估临床一致性。CSM$_{\text{CCTA}}$与放射科医生评分显示出强相关性(皮尔逊相关系数$r=0.97$,$p<0.001$),优于次优指标FORTE($r=0.70$)达0.27,且在160次成对比较中与专家偏好一致的有115次(71.9%)。在受控扰动下,CSM$_{\text{CCTA}}$对临床等效措辞保持稳定,并随信息逐步省略单调下降。在多中心基准中,经CCTA训练的C2RG模型在所有四家医院均取得最高CSM$_{\text{CCTA}}$评分,尽管其性能仍远未达到最优。相比之下,通用模型的输出中与CCTA无关的报告占比高达98.7%。总体而言,该基准为模型比较提供了标准化设置,而CSM$_{\text{CCTA}}$支持对发现一致性和解剖学特异性进行临床结构化评估,这些结果为评估CCTA报告生成提供了更贴合临床、更具解剖学分辨率的方法。代码可在该https URL获取。
英文摘要
Reliable evaluation of automated coronary computed tomography angiography (CCTA) report generation requires standardized multicentre benchmarks and clinically structured metrics. We established a four-centre benchmark comprising 3,021 CCTA series from 818 patient-report pairs to evaluate seven open-source three-dimensional vision-language models. We developed CSM$_{\text{CCTA}}$, a clinically structured metric for CCTA report evaluation, with patient-, vessel-, and segment-level variables defined according to clinical guidelines. Report pairs are compared at the finest shared anatomical level, and the contributions of different clinical components are weighted based on expert assessments. We estimated these weights using 70 expert-scored cases and evaluated clinical alignment in a non-overlapping set of 30 cases. CSM$_{\text{CCTA}}$ showed a strong correlation with radiologist scores (Pearson's $r=0.97$, $p<0.001$), exceeding the next-best metric, FORTE ($r=0.70$), by 0.27, and agreed with expert preferences in 115 of 160 pairwise comparisons (71.9\%). Under controlled perturbations, CSM$_{\text{CCTA}}$ remained stable to clinically equivalent wording and decreased monotonically with progressive information omission. In the multicenter benchmark, the CCTA-trained C2RG model achieved the highest CSM$_{\text{CCTA}}$ scores across all four hospitals, although its performance remained far from optimal. In contrast, CCTA-irrelevant reports accounted for up to 98.7\% of the outputs from generalist models. Together, the benchmark provides a standardized setting for model comparison, while CSM$_{\text{CCTA}}$ enables clinically structured evaluation of finding agreement and anatomical specificity. These results support a more clinically aligned and anatomically resolved approach to evaluating CCTA report generation. Code is available at https://openi.pcl.ac.cn/OpenMedIA/CSM_CCTA.