arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

DualDiT:用于联合生成OCT图像与分割掩码的条件双输出扩散Transformer

DualDiT: A Conditional Dual-Output Diffusion Transformer for Joint OCT Image and Segmentation Mask Generation

Fernando García-Torres, Rocío del Amor, Sandra Morales, Álvaro Barroso, Peter Heiduschka, Björn Kemper, Valery Naranjo

arXiv 2607.29337首次发表:更新:

发表机构

Instituto Universitario de Investigación en Tecnología Centrada en el Ser Humano (HUMAN-tech), Universitat Politècnica de València; Artikode Intelligence S.L; Biomedical Technology Center of the Medical Faculty, University of Muenster; Department of Ophthalmology, University of Muenster Medical Centre(瓦伦西亚理工大学以人类为中心技术大学研究所; Artikode智能公司; 明斯特大学医学院生物医学技术中心; 明斯特大学医学中心眼科)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出DualDiT模型,联合生成小鼠OCT图像与分割掩码,在生成质量、下游分割性能上优于DDPM、LDM,可用于医学图像数据增强

AI 中文摘要

背景与目标:生成具有解剖学准确分割掩码的真实医学图像,有助于解决医学成像中带标注数据短缺的问题,尤其是小鼠眼的光学相干断层扫描(OCT)领域,由于视网膜结构微小且需要专业知识,手动描绘视网膜层十分费力,导致数据集稀缺。尽管扩散模型在医学图像合成中表现良好,但联合图像-掩码生成主要依赖基于U-Net的去噪器,扩散Transformer在很大程度上未被探索。方法:我们提出一种条件双输出扩散Transformer(DualDiT),用于联合生成离体小鼠视网膜的OCT B扫描图像和上层视网膜细胞层的分割掩码。DualDiT通过预训练VAE将两种模态编码到共享潜在空间,拼接它们的潜在表示,并对联合张量执行条件扩散。我们将DualDiT与两种适配的扩散基线DDPM和LDM进行比较,通过Fréchet Inception Distance(FID)和空间FID(sFID)评估生成质量;通过合成数据增强下游U-Net分割任务评估实用效用;并通过三名领域专家评估感知真实性。结果:DualDiT取得最佳生成质量(FID 56.14,sFID 114.35),优于DDPM和LDM;专家小组将46%的合成样本误判为真实样本,42%的真实样本误判为合成样本;添加DualDiT生成的图像和掩码提升了保留的分割测试集上的Dice和IoU分数。结论:DualDiT表明基于Transformer的扩散模型可有效学习OCT图像与分割掩码的联合分布,在生成保真度、下游效用和感知真实性方面优于基于DDPM和LDM的基线,凸显其在标注稀缺的医学成像中用于数据增强的潜力。

英文摘要

Background and Objective: Generating realistic medical images with anatomically accurate segmentation masks helps address the shortage of annotated data in medical imaging, particularly in optical coherence tomography (OCT) of mouse eyes, where manual retinal layer delineation is labour-intensive due to tiny structures and required expertise, resulting in scarce datasets. While diffusion models perform well in medical image synthesis, joint image-mask generation has relied mainly on U-Net-based denoisers, leaving diffusion transformers largely unexplored. Methods: We propose a conditional dual-output Diffusion Transformer (DualDiT) for joint synthesis of OCT B-scans and segmentation masks of the upper retinal cell layers in ex vivo mouse retina. DualDiT encodes both modalities into a shared latent space via a pretrained VAE, concatenates their latent representations, and performs conditional diffusion over the joint tensor. We compared DualDiT against two adapted diffusion baselines: DDPM and LDM. Generative quality was assessed via Fréchet Inception Distance (FID) and spatial FID (sFID); practical utility via synthetic data augmentation for downstream U-Net segmentation; and perceptual realism via evaluation by three domain experts. Results: DualDiT achieved the best generative quality (FID 56.14, sFID 114.35), outperforming DDPM and LDM. Expert panels misclassified 46% of synthetic samples as real and 42% of real samples as synthetic. Adding DualDiT-generated images and masks improved Dice and IoU scores on a held-out segmentation test set. Conclusions: DualDiT shows that transformer-based diffusion models can effectively learn the joint distribution of OCT images and segmentation masks, surpassing DDPM- and LDM-based baselines in generative fidelity, downstream utility, and perceptual realism, highlighting its potential for data augmentation in annotation-scarce medical imaging.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑