发表机构
Università Campus Bio-Medico di Roma; Umeå University; UniCamillus – Saint Camillus International University of Health Sciences(罗马 Campus Bio-Medico 大学; 于默奥大学; UniCamillus – 圣卡米勒斯国际健康科学大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出基于全容积多任务潜在流匹配的方法,解耦容积先验与跨模态映射目标,训练单个多任务模型实现跨/模态内转换,在全容积处理、零样本泛化等方面优于基线,为合成系统泛化提供可扩展途径。
AI 中文摘要
跨模态医学图像转换可减轻多模态采集的负担,但该领域仍受两个相互关联的限制约束:现有方法针对2D切片或3D块而非全容积操作,且为每个转换任务训练单独的模型。这两个问题源于单一原因,即缺乏足够强的容积先验,这迫使生成模型同时学习解剖外观和跨模态映射,在现有配对数据集规模下这是一个不适定问题。我们建议将这些目标解耦。一个大规模预训练的3D变分自编码器提供了容积外观的紧凑潜在表示,将转换简化为条件流匹配问题。这种压缩使全容积处理变得可行,同时分辨率感知采样策略保留了原生解剖尺度。我们在三个多中心数据集上联合训练单个模型,涵盖跨模态(MRI→CT、CBCT→CT)和模态内(MRI→MRI)任务。在所有任务中,全容积处理优于基于块的对应方法,且多任务模型与特定任务基线性能相当,同时将N个网络替换为1个网络。关键的是,联合训练解锁了特定任务方法无法实现的能力:对训练期间未见过的解剖区域的零样本泛化,与全监督模型的SSIM差距在0.15以内,以及沿从未直接监督的路径进行的组合跨数据集转换。这些结果表明,将强大的容积先验与多任务训练相结合,是构建能泛化到训练分布之外的合成系统的可扩展途径。代码可在此https URL获取。
英文摘要
Cross-modality medical image translation can reduce the burden of multi-modal acquisitions, yet the field remains constrained by two coupled limitations: methods operate on 2D slices or 3D patches rather than whole volumes, and train a separate model for each translation task. Both stem from a single cause, the absence of a sufficiently strong volumetric prior, which forces generative models to learn anatomical appearance and cross-modality mapping simultaneously, an ill-posed problem at the scale of available paired datasets. We propose to decouple these objectives. A large-scale pretrained 3D variational autoencoder provides a compact latent representation of volumetric appearance, reducing translation to a conditional flow-matching problem. This compression makes whole-volume processing tractable, while a resolution-aware sampling strategy preserves native anatomical scale. We train a single model jointly across inter-modality (MRI$\to$CT, CBCT$\to$CT) and intra-modality (MRI$\to$MRI) tasks over three multi-center datasets. Across all tasks, whole-volume processing outperforms its patch-based counterpart, and the multi-task model matches task-specific baselines while replacing $N$ networks with one. Crucially, joint training unlocks capabilities inaccessible to task-specific approaches: zero-shot generalization to anatomical regions unseen during training, within 0.15 SSIM of the fully supervised model, and compositional cross-dataset translation along paths never directly supervised. These results suggest that combining a strong volumetric prior with multitask training is a scalable route toward synthesis systems that generalize beyond their training distribution. Code is available at https://github.com/arco-group/Whole-Volume-Latent-FM.
CommentsAccepted ad Sashimi 2026