用于合成组织病理学图像生成的条件扩散模型评估
Assessment of Conditional Diffusion Model for Synthetic Histopathology Image Generation
浏览论文内容
中文总结 AI 辅助
本研究针对合成组织病理学图像生成的评估问题,提出基于数字病理学预训练基础模型改进的FID、IS及精确率-召回率指标,实验发现改进后的IS与下游细胞核分割性能相关性更高,且生成数据多样性对分割性能的提升作用强于单张图像视觉保真度。
中文摘要 AI 辅助
合成组织病理学图像生成已成为一种可解决计算病理学中数据稀缺问题的方法,但当前的评估方法可能无法充分评估医学应用中合成数据的质量。本研究调查并解决了现有评估指标的局限性,探索了一种通过领域特定指标和下游任务验证来评估合成组织病理学图像质量的方法。研究表明,诸如Frechet Inception Distance(FID)和Inception Score(IS)之类的传统合成数据评估指标在应用于组织病理学图像时可能存在局限性,因为它们依赖于在ImageNet上预训练的特征提取器。为解决这些局限性,我们提出了使用在数字病理学数据集上预训练的基础模型改进的FID和IS方法,并补充了基于精确率-召回率的指标作为额外质量评估的一部分。我们使用在四个基准数据集上训练的条件去噪扩散模型,采用两步训练方法,生成了具有系统变化质量特征的合成数据集。我们还使用包括aggregated Jaccard index(AJI+)和Dice系数在内的常用指标,测量了合成数据质量指标与下游细胞核分割性能之间的相关性。研究结果表明,病理学特定指标可能提供更好的判别能力。具体而言,改进后的Inception Score与下游任务性能的相关性更高(与AJI+的相关系数r=0.6096,p=0.0122),而原始IS与AJI+的相关系数r=0.0708,p=0.7944。我们的观察结果表明,与提高单个生成图像的视觉保真度相比,增加生成训练数据的多样性与分割模型性能具有更高的正相关性。
英文摘要
Synthetic histopathology image generation has emerged as an approach that may address data scarcity in computational pathology, yet current evaluation methodologies may not fully assess synthetic data quality for medical applications. This work investigates and addresses limitations in existing evaluation metrics, investigating an approach for assessing synthetic histopathology image quality through domain-specific metrics and downstream task validation. We show that conventional synthetic data evaluation metrics such as Frechet Inception Distance (FID) and Inception Score (IS) may have limitations when applied to histopathology images due to their reliance on ImageNet-pretrained feature extractors. To address these limitations, we propose for consideration modified FID and IS approaches utilizing foundation models pretrained on digital pathology datasets, supplemented by precision-recall based metrics as part of an additional quality assessment. Using conditional denoising diffusion models trained on four benchmark datasets, with a two-step training approach, we generated synthetic datasets with systematically varied quality characteristics. We also measured the correlation between the synthetic data quality metrics with downstream nuclei segmentation performance using common metrics including the aggregated Jaccard index (AJI+) and the Dice coefficient. The study results suggest that pathology-specific metrics may provide improved discriminative power. Specifically, the modified Inception Score indicates higher correlation with downstream task performance (r=0.6096 with AJI+, p=0.0122), compared to the original IS (r=0.0708, p=0.7944). Our observations indicate that increasing the variety of generated training data has a higher positive correlation with segmentation model performance than improving the visual fidelity of individual generated images.