CARDEA:基于空间证据的可审计推理用于冠状动脉造影端到端解读
CARDEA: Auditable Reasoning Grounded in Spatial Evidence for End-to-End Coronary Angiography Interpretation
浏览论文内容
中文总结 AI 辅助
CARDEA是一个基于空间证据可审计推理的端到端冠状动脉造影解读模型,通过三阶段训练(含RLVR)提升诊断准确性和零样本报告生成能力。
中文摘要 AI 辅助
有创冠状动脉造影(CAG)是诊断冠状动脉疾病的金标准,但不同观察者之间的解读差异很大。现有的人工智能系统可以提高一致性,但缺乏可审计的决策过程,并且在全面的开放式评估方面存在局限,这削弱了临床医生的信任和临床采用的准备度。我们开发了CARDEA,一个统一的大型视觉-语言模型,作为CAG流程的推理核心。它仅使用公共数据集和封闭式任务进行训练,分为三个阶段:视觉特征对齐、自蒸馏的Chain-of-Box(CoB)冷启动,以及带有可验证奖励的强化学习(RLVR),其中CoB奖励鼓励在推理轨迹中使用边界框。我们将其两项研究级诊断——优势分类和复杂性评估——与专用分类器和两位介入心脏病专家进行了比较。报告生成被排除在训练之外,并在外部队列上跨阶段进行零样本评估,使用血管严重程度宏F1分数。CARDEA在分布内优势分类上落后于分类器,但在领域偏移下与之持平(准确率0.91 [95%置信区间(CI),0.86至0.95]),在复杂性评估上与心脏病专家相当(准确率0.90 [CI,0.82至0.97])。只有RLVR改善了零样本报告生成,将其血管严重程度宏F1分数(0.686 [CI,0.664至0.707])提高到未调优基础模型(0.513)之上,并且是始终正常基线(0.312)的两倍多。CARDEA运行一个端到端的CAG流程,从原始多视角视频通过关键帧选择到研究级诊断,同时在其结论背后展示可审计的空间证据。在可验证的封闭式任务上进行RLVR,展现出了监督模仿所未能实现的开放式报告能力。临床使用需要与专家心脏病专家进行前瞻性验证。
英文摘要
Invasive coronary angiography (CAG) is the gold standard for diagnosing coronary artery disease, but interpretation varies substantially among observers. Existing AI systems can improve consistency but lack auditable decision processes and are limited in comprehensive open-ended assessment, undermining clinician trust and clinical adoption readiness. We developed CARDEA, a unified large vision-language model that serves as the inference core of a CAG pipeline. It was trained solely on public datasets and closed-ended tasks in three stages: visual feature alignment, a self-distilled Chain-of-Box (CoB) cold start, and reinforcement learning with verifiable rewards (RLVR) with a CoB reward encouraging bounding-box use in the reasoning trace. We assessed its two study-level diagnoses, dominance classification and complexity assessment, against a dedicated classifier and two interventional cardiologists. Report generation was excluded from training and evaluated zero-shot across stages on an external cohort using vessel-severity macro-$F_1$. CARDEA trailed the classifier on in-distribution dominance but drew level under domain shift (accuracy, 0.91 [95% confidence interval (CI), 0.86 to 0.95]) and was comparable to the cardiologists on complexity assessment (accuracy, 0.90 [CI, 0.82 to 0.97]). Only RLVR improved zero-shot report generation, raising its vessel-severity macro-$F_1$ (0.686 [CI, 0.664 to 0.707]) above the untuned base model (0.513) and over twice the always-normal floor (0.312). CARDEA runs an end-to-end CAG pipeline from raw multi-view videos through keyframe selection to study-level diagnosis while exposing auditable spatial evidence behind its conclusions. RLVR on verifiable closed-ended tasks surfaced open-ended reporting ability that supervised imitation did not. Clinical use requires prospective validation against expert cardiologists.
发表机构
- Artificial Intelligence and Robotics Innovation Center(人工智能与机器人创新中心)
- China Medical University Hospital(中国医药大学附设医院)
- China Medical University(中国医药大学)
- Neuroscience and Brain Disease Center(神经科学与脑疾病中心)
- School of Medicine(医学院)
机构由 AI 辅助整理,请以论文原文为准。