arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.16491physics.med-ph

用于医学生成与重建的连续3D潜扩散

Continuous 3-D Latent Diffusion for Medical Image Generation and Reconstruction

Youness Mellak, Antoine De Paepe, Dimitris Visvikis, Alexandre Bousse

首次发表
浏览论文内容

中文总结 AI 辅助

研究针对高分辨率三维医学扩散模型成本问题,提出连续3D潜扩散模型框架,核心是含特定解码器的紧凑自动编码器,经实验评估,该框架在计算效率、内存使用和重建效果间达平衡,能支持多种医学影像任务。

中文摘要 AI 辅助

高分辨率三维医学扩散模型受全容积处理成本限制,即便在紧凑潜空间去噪也是如此。我们引入了一种用于计算机断层扫描(CT)和磁共振成像(MRI)生成及测量引导重建的连续3D潜扩散模型(LDM)框架。其核心组件是一个紧凑自动编码器(AE),带有坐标条件局部隐式图像函数(LIIF)解码器,将容积表示为空间坐标的连续函数。通过在潜网格上对卷积解码器评估一次,并将重复计算限制在轻量级隐式头部,该设计避免了重叠子容积解码,同时对于逆问题优化保持可微。我们在512^3体素的CT容积和256^3体素的MRI容积上评估了该框架。在高分辨率CT上,所提出的AE比评估的参考自动编码器快约12 - 32倍,GPU内存使用最低,尽管体素级精度略有降低,但保留了相当的结构保真度。由此产生的冻结3D潜先验生成连贯的全容积且无可见补丁接缝,可通过硬数据一致性应用于稀疏视图CT和加速MRI重建,无需特定任务再训练。虽然直接像素域重建更准确,但结果表明单个容积潜先验可在一个GPU上支持无条件生成和测量条件重建。总体而言,该框架在连续容积解码、计算效率和细节保留之间提供了实际的权衡。我们的代码将在这个https URL上提供。

英文摘要

High-resolution three-dimensional (3-D) medical diffusion models remain constrained by the cost of processing full volumes, even when denoising is performed in a compact latent space. We introduce a continuous 3-D latent diffusion model (LDM) framework for computed tomography (CT) and magnetic resonance imaging (MRI) generation and measurement-guided reconstruction. Its central component is a compact autoencoder (AE) with a coordinate-conditioned local implicit image function (LIIF) decoder that represents a volume as a continuous function of spatial coordinates. By evaluating the convolutional decoder once on the latent grid and restricting repeated computation to a lightweight implicit head, the proposed design avoids overlapping sub-volume decoding while remaining differentiable for inverse-problem optimization. We evaluate the framework on CT volumes of 512^3 voxels and MRI volumes of 256^3 voxels. On high-resolution CT, the proposed AE is approximately x12-32 faster than the evaluated reference autoencoders, achieves the lowest peak graphics processing unit (GPU) memory use, and retains comparable structural fidelity despite a moderate reduction in voxel-level accuracy. The resulting frozen 3-D latent prior generates coherent full volumes without visible patch seams and can be applied, without task-specific retraining, to sparse-view CT and accelerated MRI reconstruction through hard data consistency. Although direct pixel-domain reconstruction remains more accurate, the results demonstrate that a single volumetric latent prior can support both unconditional generation and measurement-conditioned reconstruction on one GPU. Overall, the framework provides a practical trade-off between continuous volumetric decoding, computational efficiency, and fine-detail preservation. Our code will be made available at https://github.com/mellak/.

补充信息

↑