发表机构
The University of Melbourne(墨尔本大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究提出TCAM-Diff这一3D医学图像生成模型,采用仅解码器自动编码器方法学习三平面表示,利用三平面感知交叉注意力扩散模型整合特征,经预训练解码器模块转换为3D体素,在多规模医学数据集实验中结果出色,优于其他类似方法。
AI 中文摘要
我们介绍了TCAM-Diff,一种新颖的3D医学图像生成模型,它降低了编码和生成高分辨率3D数据的内存需求。该模型采用仅解码器的自动编码器方法从密集体素中学习三平面表示,并利用泛化操作防止过拟合。随后,它使用三平面感知交叉注意力扩散模型有效学习和整合这些特征。此外,扩散模型生成的特征可通过预训练的解码器模块快速转换为3D体素。我们在BrainTumour 128 x 128 x 128、Pancreas 256 x 256 x 2穿骸扁缴壮剂憋烯铂楼56和Colon 512 x 512 x 512这三种不同规模的医学数据集上进行实验,结果出色。我们用MSE和SSIM评估重建质量,利用Wasserstein生成对抗网络(W-GAN)评判器评估生成质量。与现有方法比较表明,我们的方法比具有相似大小潜在空间的其他编码器-解码器方法有更好的重建和生成结果。
英文摘要
We introduce TCAM-Diff, a novel 3D medical image generation model that reduces the memory requirements to encode and generate high-resolution 3D data. This model utilizes a decoder-only autoencoder method to learn triplane representation from dense volume and leverages generalization operations to prevent overfitting. Subsequently, it uses a triplane-aware cross-attention diffusion model to learn and integrate these features effectively. Furthermore, the features generated by the diffusion model can be rapidly transformed into 3D volumes using a pre-trained decoder module. Our experiments on three different scales of medical datasets, BrainTumour 128 x 128 x 128, Pancreas 256 x 256 x 256, and Colon 512 x 512 x 512, demonstrate outstanding results. We utilized MSE and SSIM to assess reconstruction quality and leveraged the Wasserstein Generative Adversarial Network (W-GAN) critic to assess generative quality. Comparisons with existing approaches show that our method gives better reconstruction and generation results than other encoder-decoder methods with similar-sized latent spaces.
CommentsAccepted at AAAI 2025. Code is available at https://github.com/Fredy-Zhang/TCAM-Diff
Journal refProceedings of the AAAI Conference on Artificial Intelligence, 39(21): 22732-22740, 2025