arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.08109cs.CV

超越从零训练:用于数据高效且可泛化的心脏MRI重建的基础模型

Beyond Training from Scratch: Foundation Models for Data-Efficient and Generalizable Cardiac MRI Reconstruction

Anam Hashmi, Mayug Maniparambil, Julia Dietlmeier, Kathleen M. Curran, Noel E. O'Connor

首次发表
浏览论文内容

中文总结 AI 辅助

本研究提出利用预训练视觉基础模型(如CLIP、BiomedCLIP、DINOv2)作为心脏MRI重建的先验,在CMRxRecon基准上相比从零训练的Transformer展现出更好的数据效率和泛化能力,其中DINOv2表现最优。

中文摘要 AI 辅助

心脏磁共振成像重建旨在从欠采样采集中恢复高质量图像,从而实现更快的扫描同时保持诊断保真度。近年来的重建方法通常从零开始训练,并且往往需要大量特定任务的数据,这限制了它们在数据稀缺和分布偏移情况下的鲁棒性。在这项工作中,我们研究了预训练的视觉基础模型是否可以作为加速心脏MRI重建的有效先验。我们提出了一种重建框架,该框架在基于Transformer的重建架构中集成了冻结的和参数高效适配的视觉编码器,包括CLIP、BiomedCLIP和DINOv2。在CMRxRecon2023和CMRxRecon2024基准上的大量实验表明,预训练表示在多个加速因子下始终优于从零训练的Transformer。我们进一步评估了在有限监督和跨数据集迁移下的性能,表明基础模型提供了优越的数据效率和泛化能力。虽然冻结表示在极端低数据场景中特别有效,但当有中等数量的训练数据可用时,低秩适配(LoRA)带来了额外的增益。在评估的骨干网络中,DINOv2实现了最强的整体性能。这些发现凸显了视觉基础模型作为心脏MRI重建的鲁棒且可迁移先验的潜力。

英文摘要

Cardiac magnetic resonance imaging reconstruction aims to recover high-quality images from undersampled acquisitions, enabling faster scans while preserving diagnostic fidelity. Recent reconstruction methods are typically trained from scratch and often require large amounts of task-specific data, limiting their robustness under data scarcity and distribution shifts. In this work, we investigate whether pretrained vision foundation models can serve as effective priors for accelerated cardiac MRI reconstruction. We propose a reconstruction framework that integrates frozen and parameter-efficiently adapted visual encoders, including CLIP, BiomedCLIP, and DINOv2, within a transformer-based reconstruction architecture. Extensive experiments on the CMRxRecon2023 and CMRxRecon2024 benchmarks demonstrate that pretrained representations consistently outperform a transformer trained from scratch across multiple acceleration factors. We further evaluate performance under limited supervision and cross-dataset transfer, showing that foundation models provide superior data efficiency and generalization. While frozen representations are particularly effective in extreme low-data regimes, Low-Rank Adaptation (LoRA) yields additional gains when moderate amounts of training data are available. Among the evaluated backbones, DINOv2 achieves the strongest overall performance. These findings highlight the potential of vision foundation models as robust and transferable priors for cardiac MRI reconstruction.

发表机构

  • Research Ireland Centre for Research Training in Machine Learning(爱尔兰研究机器学习研究中心)
  • Dublin City University(都柏林城市大学)
  • University College Dublin(都柏林大学学院)
  • Rinn Artificial Intelligence(Rinn人工智能公司)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑