发表机构
University of Rochester Medical Center; Biomedical Engineering, RIT(罗切斯特大学医学中心; 罗切斯特理工大学生物医学工程系)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出两阶段VLM框架用于左心房LGE-MRI的临床导向图像质量评估,在60个图像切片-文本对数据集上,InternVL2的标准级准确率最高,DeepSeek实现完美临床可用性一致性。
AI 中文摘要
LGE心脏磁共振成像(LGE cardiac MRI)广泛用于心房颤动患者的左心房纤维化评估及消融规划,因为从LGE-MRI中识别出的纤维化组织区域信息对导管消融至关重要。消融规划期间使用的图像质量不佳常导致消融靶点定位错误,直接影响手术安全性和结局。目前,扫描是否达到消融规划最低质量阈值的决定由阅片放射科医生非正式做出,未被任何自动化系统捕获,但这无疑是图像质量评估(IQA)过程中对安全性最关键的输出。然而,噪声、运动伪影和边界清晰度差导致的图像质量变化会严重损害下游分割和临床决策任务的可靠性。放射科专家的人工质量评估具有主观性且难以规模化,而现有自动化方法仅生成标量分数,缺乏可解释的临床推理。本研究提出一种两阶段视觉语言模型(VLM)框架,用于左心房LGE-MRI的临床导向图像质量评估。第一阶段,微调后的VLM生成结构化放射科风格质量报告,预测放射科医生定义的五项标准:噪声、运动伪影、左心房(LA)边界准确性、肺静脉(PV)区域准确性及欠分割严重程度。第二阶段,基于GPT的推理模块将预测的质量和报告映射为结构化质量分数及用于消融规划的二元临床可用性决策。我们整理了包含20名患者的60个带注释图像切片-文本对的数据集,并对四种最先进的VLM架构进行基准测试。InternVL2达到最高的标准级准确率(平均准确率=0.65,皮尔逊线性相关系数=0.79),而DeepSeek实现了完美的临床可用性一致性(准确率=1.00,科恩kappa系数=1.00)。
英文摘要
LGE cardiac MRI is widely used for left atrial fibrosis assessment and ablation planning in atrial fibrillation patients as knowledge of fibrotic tissue regions identified from LGE-MRI is critical for catheter ablation. Often, poor quality images used during ablation planning can cause mis-localization of ablation targets, directly impacting procedure safety and outcome. The decision of whether a scan meets the minimum quality threshold for ablation planning is currently made informally by the reviewing radiologist and is not captured by any automated system, yet it is arguably the most safety-critical output of the image quality assessment (IQA) process. However, variations in image quality caused by noise, motion artifacts, and poor boundary definition significantly compromise the reliability of downstream segmentation and clinical decision-making tasks. Manual quality assessment by expert radiologists is subjective and difficult to scale, while existing automated methods produce scalar scores without interpretable clinical reasoning. In this work, we propose a two-stage vision language model (VLM) framework for clinically grounded image quality assessment of left atrial LGE-MRI. In the first stage, a fine-tuned VLM generates structured radiology-style quality reports predicting five radiologist-defined criteria: Noise, Motion Artifact, LA Boundary Accuracy, PV Region Accuracy, and Under-segmentation Severity. In the second stage, a GPT-based reasoning module maps the predicted quality and reports to a structured quality scores and binary clinical usability decision for ablation planning. We curate a dataset of 60 annotated image slice-text pairs from 20 patients and benchmark four state-of-the-art VLM architectures. InternVL2 achieves the highest criterion-level accuracy (Avg ACC=0.65, PLCC=0.79), while DeepSeek achieves perfect clinical usability agreement (Acc=1.00, kappa=1.00).