arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

面向条件化CMR合成的生成式基础模型的元数据感知适配

Metadata-Aware Adaptation of a Generative Foundation Model for Conditional CMR Synthesis

Marc Rodríguez, Grzegorz Skorupko, Nay Aung, Steffen E Petersen, Karim Lekadir, Polyxeni Gkontra

arXiv 2608.24342首次发表:更新:

发表机构

Facultat de Matemàtiques i Informàtica, Universitat de Barcelona; NIHR Barts Biomedical Research Centre; St Bartholomew’s Hospital; Institució Catalana de Recerca i Estudis Avançats (ICREA)(巴塞罗那大学数学与信息学院; 英国国立卫生研究院巴特生物医学研究中心; 圣巴塞洛缪医院; 加泰罗尼亚高级研究学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出整合无元数据CFG等策略的元数据感知适配方法,基于预训练潜在扩散模型实现仅用患者元数据的CMR合成,在UK Biobank数据集上显著提升FID,验证了生成式基础模型用于临床CMR合成的潜力。

AI 中文摘要

合成图像生成是解决医学成像中数据稀缺及临床重要表型代表性不足问题的有前景策略,但生成能忠实反映患者有意义特征的图像仍具挑战性。本研究使用预训练的潜在扩散模型探究元数据条件化心脏磁共振(CMR)合成,将结构化临床元数据和切片位置编码为文本提示以引导CMR生成。为提升元数据依从性并解决临床属性不平衡问题,我们整合了三种策略:无元数据的无分类器引导(CFG)、对比批次处理、逆频率采样。该框架在英国生物银行(UK Biobank)的59058个短轴CMR上进行微调与评估,采用配对图像相似性、分布保真度及亚组水平分析。该组合方法的Fréchet Inception Distance(FID)为37.47,较未使用这些策略微调的同一模型提升57.04%,较需心脏几何结构作为额外输入的现有文本条件化CMR扩散基线提升28.68%,且仅依赖患者元数据。这种主要由无元数据CFG驱动的分布增益伴随配对相似性的小幅降低,表明模型优先考虑群体层面的真实性而非精确图像复现。亚组分析显示,人口统计学与采集相关元数据的对齐性得到改善,而疾病特异性条件化是最具挑战性的任务。这些发现证明了生成式基础模型在临床有意义CMR合成方面的潜力,同时凸显了对更有效元数据感知条件化策略的需求。我们的代码可在此URL获取。

英文摘要

Synthetic image generation is a promising strategy to address data scarcity and the underrepresentation of clinically important phenotypes in medical imaging, yet generating images that faithfully reflect meaningful patient characteristics remains challenging. In this work, we investigate metadata-conditioned cardiac magnetic resonance (CMR) synthesis using a pretrained latent diffusion model, encoding structured clinical metadata and slice position as textual prompts to guide CMR generation. To improve metadata adherence and address the imbalance of clinical attributes, we integrate three strategies: Metadata-Free Classifier-Free Guidance (CFG), Contrastive Batching, and Inverse-Frequency Sampling. The framework was fine-tuned and evaluated on 59,058 short-axis CMR from the UK Biobank using paired image similarity, distributional fidelity, and subgroup-level analyses. The combined approach achieved a Fréchet Inception Distance (FID) of 37.47, improving by 57.04\% over the same model fine-tuned without these strategies and by 28.68\% over a previous text-conditioned CMR diffusion baseline requiring cardiac geometry as additional input, while relying solely on patient metadata. This distributional gain, driven mainly by Metadata-Free CFG, came with a modest reduction in paired similarity, suggesting that the model prioritizes population-level realism over exact image reproduction. Subgroup analyses demonstrated improved alignment across demographic and acquisition-related metadata, with disease-specific conditioning being the most challenging task. These findings demonstrate the potential of generative foundation models for clinically meaningful CMR synthesis while highlighting the need for more effective metadata-aware conditioning strategies. Our code is available at https://github.com/rodriguezmarc/conditional-cmr.

Comments5 pages, 3 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑