arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.13455cs.CV

基于基础模型潜在空间的临床可操控视网膜图像生成评估

Evaluation of Clinically Steerable Retinal Image Generation from Foundation Model Latent Spaces

Zuzanna A. Wakefield-Skórniewska, Bartłomiej W. Papież

首次发表
浏览论文内容

中文总结 AI 辅助

该研究评估了四种视网膜基础模型的可控图像生成能力,发现其在自身框架内生成的视网膜图像优于传统潜在扩散,但与真实图像存在表征差距,需进一步对齐。

中文摘要 AI 辅助

医学基础模型学习具有临床意义表型的潜在表征,但其支持可控图像生成的能力在很大程度上仍未被探索。我们在表征分词器框架内评估了四种视网膜基础模型,研究基础模型潜在表征中编码的人口统计学和临床信息是否在合成图像生成过程中得到保留。结果显示,在其所属的基础模型内评估时,生成的表征和图像忠实地继承了表型信息,在多个下游预测任务中始终优于传统潜在扩散模型。然而,使用在真实图像上训练的分类器进行评估时,这些优势基本消失,这揭示了一种此前未被表征的合成到真实的表征差距。这些发现表明,基础模型潜在空间为可控视网膜合成提供了强大的基础,同时凸显了需要更好地使合成表征与真实图像分布对齐。

英文摘要

Medical foundation models learn latent representations of clinically meaningful phenotypes, yet their ability to support controllable image generation remains largely unexplored. We evaluate four retinal foundation models within the representation tokenizer framework and examine whether demographic and clinical information encoded in latent representations from foundation models is preserved during synthetic image generation. We show that generated representations and images faithfully inherit phenotype information when evaluated within their originating foundation models, consistently outperforming conventional latent diffusion on multiple downstream prediction tasks. However, these gains largely disappear when evaluated using classifiers trained on real images, revealing a previously uncharacterised synthetic-to-real representation gap. These findings demonstrate that foundation-model latent spaces provide a powerful substrate for controllable retinal synthesis while highlighting the need to better align synthetic representations with real-image distributions.

发表机构

  • Nuffield Department of Population Health, University of Oxford(牛津大学纳菲尔德人口健康系)
  • University of Oxford(牛津大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑