arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.08849eess.IVcs.CV

基于语义标签的腹部超声模拟:配对标签到物理图像翻译

Abdominal Ultrasound Simulation from Semantic Labels using Paired Label-to-Physics-Based Image Translation

Santiago Vitale, Duilio Deangeli, Ignacio Larrabide, José Ignacio Orlando

首次发表
浏览论文内容

中文总结 AI 辅助

提出两阶段学习流程,从语义标签生成腹部超声图像,无需推理时CT,支持可控变形与病理模拟,SDM优于Pix2Pix。

中文摘要 AI 辅助

目的:当前的腹部超声(US)模拟方法通常需要基于CT的解剖参考进行射线投射,这限制了变形和病理变异的可能性。我们提出了一种基于学习的流程,该流程被训练用于从语义标签预测由CT扫描导出的物理图像,从而在推理时无需患者特异性CT体积即可实现受控模拟。方法:我们引入了一个两阶段流程,通过一个简化的超声图像将解剖分割映射到逼真的超声图像。第一阶段利用基于CT射线投射输出训练的模型,从语义标签合成该图像。第二阶段利用解剖引导的非配对翻译将其细化为逼真的超声扫描。变形和病理通过编辑解剖图谱生成。结果:我们在第一阶段评估了Pix2Pix和语义扩散模型(SDM),随后在第二阶段使用分割引导的CycleGAN(SG-CycleGAN)进行细化。SDM在形态学指标上显著优于Pix2Pix,包括MAE(19.85对21.65)、SSIM(0.28对0.24)和mIoU(0.43对0.29),而Pix2Pix在感知点估计上表现更好(LPIPS:0.17对0.19;FID:0.32对0.37;KID:0.25对0.48)。结论:使用基于物理的监督训练配对生成模型,能够在推理时直接从语义标签近似CT导出的射线投射输出。尽管该流程在推理时不需要患者特异性CT体积,但CT导出的分割和射线投射模拟仍然是训练第一阶段所必需的。一旦训练完成,该框架能够通过语义图谱修改实现可控的健康和病理模拟。

英文摘要

Purpose: Current abdominal ultrasound (US) simulation methods often require CT-based anatomical references for ray-casting, limiting deformation and pathology variability. We propose a learning-based pipeline trained to predict physics-based images derived from CT scans from semantic labels, enabling controlled simulations without patient-specific CT volumes at inference time. Methods: We introduce a two-stage pipeline that maps anatomical segmentations to realistic US images through a simplified US image. Stage~I synthesizes this image from semantic labels using models trained on CT-based ray-casting outputs. Stage~II refines it into a realistic US scan using anatomically guided unpaired translation. Deformations and pathologies are generated by editing anatomical maps. Results: We evaluated Pix2Pix and the Semantic Diffusion Model (SDM) in Stage~I, followed by segmentation-guided CycleGAN (SG-CycleGAN) refinement in Stage~II. SDM significantly outperformed Pix2Pix in morphological metrics, including MAE (19.85 vs. 21.65), SSIM (0.28 vs. 0.24), and mIoU (0.43 vs. 0.29), whereas Pix2Pix yielded better perceptual point estimates (LPIPS: 0.17 vs. 0.19; FID: 0.32 vs. 0.37; KID: 0.25 vs. 0.48). Conclusion: Training paired generative models with physics-based supervision enables approximation of CT-derived ray-casting outputs at inference time directly from semantic labels. Although the pipeline does not require patient-specific CT volumes at inference time, CT-derived segmentations and ray-casting simulations remain necessary to train Stage~I. Once trained, the framework enables controllable healthy and pathological simulations through semantic-map modification.

发表机构

  • National Scientific and Technical Research Council (CONICET)(国家科学技术研究委员会(CONICET))
  • UNICEN(国立中央布宜诺斯艾利斯省大学)

机构由 AI 辅助整理,请以论文原文为准。

↑