arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2506.00633cs.CVcs.AI

从对齐到合成:用于文本到CT生成的对比体 grounding

From Alignment to Synthesis: Contrastive Volumetric Grounding for Text-to-CT Generation

  • Unit of Artificial Intelligence and Computer Systems, Department of Engineering, Università Campus Bio-Medico di Roma(人工智能与计算机系统单位,工程系,罗马生物医学大学)
  • Department of Diagnostics and Intervention, Biomedical Engineering and Radiation Physics, Umeå University(诊断与干预部门,生物医学工程与放射物理学,乌梅拉大学)

机构由 AI 辅助整理,请以论文原文为准。

Daniele Molino, Camillo Maria Caruso, Filippo Ruffini, Paolo Soda, Valerio Guarrasi

AI总结:

该研究针对文本到CT生成的视觉-语言对齐瓶颈,提出带结构化难负样本的3D-CLIP编码器,结合端到端潜在扩散模型,在CT-RATE数据集上实现了优于现有方法的性能。

AI中文摘要:

从放射学报告生成语义可控的3D CT体积数据,不仅需要丰富的文本编码器,还需要基于体积空间的视觉-语言对齐。现有的文本到CT方法依赖仅用语言或2D视觉-语言目标预训练的编码器来条件化生成,提供的条件信号在语言表达上丰富,但在体积层面是盲目的。我们认为这是一个结构性局限:3D视觉-语言对齐的质量,而非文本编码器的丰富度,是体积扩散模型语义可控性的主要瓶颈。为解决该问题,我们提出一种面向生成的3D-CLIP编码器,该编码器采用仅在文本层面运行的结构化难负样本进行训练。此设计在不增加任何额外3D内存成本的情况下提升了对比难度,克服了体积编码器固有的小批量约束。所得编码器对完全端到端的潜在扩散模型进行条件化,该模型直接在3D潜在空间运行,消除了超分辨率流程引入的空间伪影和跨切片不一致性。通过系统的消融实验,我们建立了对齐质量与下游生成可控性之间的明确经验关联。在CT-RATE数据集的18种病理条件上进行评估,我们的方法在图像保真度和事实正确性上均达到了最先进的性能,同时比所有竞争方法需要更少的推理时间和GPU内存。代码位于此https URL。

英文摘要:

Generating semantically controllable 3D CT volumes from radiology reports requires more than a rich text encoder, it requires vision-language alignment grounded in volumetric space. Existing Text-to-CT approaches condition generation on encoders pretrained with language only or 2D vision-language objectives, providing conditioning signals that are linguistically expressive but volumetrically blind. We argue this is a structural limitation: the quality of 3D vision-language alignment, not the richness of the text encoder, is the primary bottleneck for semantic controllability in volumetric diffusion models. To address this, we propose a generation-oriented 3D-CLIP encoder trained with structured hard negatives that operate exclusively at the text level. This design increases contrastive difficulty without any additional 3D memory cost, overcoming the small-batch constraints inherent to volumetric encoders. The resulting encoder conditions a fully end-to-end latent diffusion model that operates directly in 3D latent space, eliminating the spatial artifacts and cross-slice inconsistencies introduced by super-resolution pipelines. Through systematic ablations, we establish a clear empirical link between grounding quality and downstream generative controllability. Evaluated on CT-RATE across 18 pathological conditions, our method achieves state-of-the-art performance on both image fidelity and factual correctness, while requiring less inference time and GPU memory than all competing methods. Code is at https://github.com/danielemolino/Text2CT.

补充信息

↑