发表机构
Siemens Healthineers; Technical University of Munich (TUM); TUM University Hospital; Imperial College London(西门子医疗; 慕尼黑工业大学; 慕尼黑工业大学医院; 帝国理工学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出用于心脏磁共振成像的通用视频基础模型 MR-JEPA,其基于多中心多序列无标注数据预训练,在心脏量化与疾病检测任务中性能优于对比方法,展现出临床应用潜力。
AI 中文摘要
心脏磁共振成像(CMR)可生成丰富的序列数据,如时间动态电影(cine)视频及空间钆延迟强化(LGE)/ mapping 序列栈,但多数深度学习方法仅处理单个二维切片,丢失了此类上下文信息。本文提出 MR-JEPA,一种用于 CMR 的自监督视频基础模型,它通过 tubelet 标记化、时空掩码增强,以及从二维 CMR 基础模型初始化,将 LeJEPA 扩展至三维时空输入。与仅局限于 cine 数据的现有 CMR 视频模型不同,MR-JEPA 在无标注的情况下,基于来自两个中心的 10505 名患者的多序列数据(cine、LGE、mapping)进行预训练。我们采用统一多视图门控注意力架构,在六个下游任务上评估冻结编码器:左心室射血分数(LV ejection fraction)、右心室射血分数(RV ejection fraction)、三种心肌应变(GLS、GCS、GRS)及四类疾病检测。MR-JEPA 在所有五项回归任务上均优于其他对比方法,包括一个在更多带文本监督的数据上预训练的领域特定 CMR 模型和一个自然视频基础模型,其 LV EF 的平均绝对误差(MAE)为 4.79%(相关系数 r=0.764),GLS 的 MAE 为 1.87(r=0.805),应变任务的 MAE 较基线降低 21%-27%。在疾病检测任务中,MR-JEPA 的宏平均曲线下面积(macro AUG)为 0.868,尽管采用完全自监督预训练目标,仍与领域特定基线具有竞争力。这些结果表明,统一视频编码器在临床心脏量化与诊断中,对各类 CMR 序列的稳健多视图利用具有潜力。
英文摘要
Cardiac magnetic resonance imaging (CMR) produces rich sequential data such as temporal cine videos and spatial LGE/mapping stacks, yet most deep learning approaches process individual 2D slices, discarding this context. We present MR-JEPA, a self-supervised video foundation model for CMR that extends LeJEPA to 3D spatiotemporal inputs through tubelet tokenization, spatiotemporal masking augmentation, and initialization from a 2D CMR foundation model. Unlike prior CMR video models limited to cine data, MR-JEPA is pretrained on multi-sequence data (cine, LGE, mapping) from 10,505 patients across two centers without annotations. We evaluate the frozen encoder on six downstream tasks using a unified multi-view gated attention architecture: LV ejection fraction, RV ejection fraction, three myocardial strains (GLS, GCS, GRS), and four-class disease detection. MR-JEPA outperforms other compared methods on all five regression tasks, including both a domain-specific CMR model pretrained on more data with text supervision and a natural-video foundation model, achieving an LV EF MAE of 4.79% (r =0.764) and a GLS MAE of 1.87 (r=0.805), with 21-27% MAE reductions over baselines on strain tasks. For disease detection, MR-JEPA achieved a macro AUG of 0.868, remaining competitive with the domain-specific baseline despite using a fully self-supervised pretraining objective. These results demonstrate the potential of a unified video encoder for robust, multi-view utilization of diverse CMR sequences in clinical cardiac quantification and diagnosis.
CommentsAccepted at STACOM 2026 (MICCAI 2026 peer-reviewed workshop)