arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

诊断多样性坍缩并验证掩膜条件扩散模型用于标记微管显微镜图像

Diagnosing Diversity Collapse and Validating Mask-Conditioned Diffusion for Labeled Microtubule Microscopy

Mario Koddenbrock, Frederic Rapp, Simone Reber, Erik Rodner

arXiv 2610.09957首次发表:更新:

发表机构

HTW Berlin; Berliner Hochschule für Technik (BHT); Max Planck Institute for Infection Biology(柏林工程与应用科学大学; 柏林工程技术应用科学大学; 马克斯·普朗克感染生物学研究所)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对掩膜条件扩散模型在显微镜图像生成中的多样性坍缩问题,提出领域验证特征空间中的多样性感知诊断方法,选出专家难辨真伪的检查点,并验证其下游分割性能优于公开替代方案。

AI 中文摘要

掩膜条件扩散模型已成为生物医学成像中生成标记训练数据的标准方法。但用于评估这些模型的工具是为不同问题而构建的。诸如FID和KID之类的标准指标通过ImageNet训练的特征对图像进行评分。这些特征无法迁移到显微镜图像,导致这些指标在专业领域校准不佳。即使使用领域合适的特征空间,单一的混合数字仍可能掩盖真实的失败:模型可能生成视觉上合理但迁移到下游任务效果不佳的图像。我们识别出一种特定的失败模式:在训练后期,真实感持续提升而纹理多样性坍缩。模型固定于一个颜色调色板,收敛到高相似度、低多样性的状态,而标准指标对此不予惩罚。因此,我们提出一种在领域验证的特征空间中的多样性感知诊断方法。它结合了基于与真实图像间相似性的真实感轴和基于生成样本间相似性的多样性轴。该诊断方法选出的检查点,在强制选择研究中,领域专家无法可靠地将其与真实记录区分开来。它还产生了有用的下游分割结果:仅使用合成标记数据训练的分割器,其骨架化IoU中位数与使用真实数据训练的分割器相当,且跨折方差更低,并明显优于该领域最佳的公开替代方案——参数化渲染器。我们还重现了先前工作中的数据高效超参数迁移实验,在小型标记子集上调整多个基础分割器并在真实图像上评估:使用我们的数据时迁移效果更强。我们在HuggingFace上发布了DiffuMT,包括三元组数据集、重现下游效用验证的代码,以及一个用于评估掩膜条件扩散模型的独立诊断工具。

英文摘要

Mask-conditioned diffusion models have become standard for generating labeled training data in biomedical imaging. But the tools used to evaluate them were built for a different problem. Standard metrics like FID and KID score images through ImageNet-trained features. Those features do not transfer to microscopy, leaving the metrics poorly calibrated for specialized domains. Even with a domain-appropriate feature space, a single blended number can still hide a real failure: a model can produce visually plausible images that still transfer poorly to downstream tasks. We identify a specific failure mode: Late in training, realism keeps improving while texture diversity collapses. The model settles onto a fixed color palette, converging to a high-similarity, low-diversity state that standard metrics do not penalize. Therefore, we propose a diversity-aware diagnostic in a domain-validated feature space. It combines a realism axis based on inter-similarity to real images with a diversity axis based on intra-similarity among generated samples. This diagnostic selects a checkpoint that domain experts cannot reliably distinguish from real recordings in a forced-choice study. It also yields useful downstream segmentation: a segmenter trained only on synthetically labeled data reaches a median skeletonised IoU comparable to one trained on real data, at lower cross-fold variance, and clearly ahead of the best available public alternative in this domain, a parametric renderer. We also reproduce a data-efficient hyperparameter-transfer experiment from prior work, tuning several foundation segmenters on a small labeled subset and evaluating on real images: transfer is stronger with our data. We release DiffuMT on HuggingFace, including the triplet dataset, code to reproduce the downstream-utility validation, and a standalone diagnostic tool for evaluating mask-conditioned diffusion models.

CommentsAccepted at NLDL 2027 (Spotlight) https://huggingface.co/spaces/HTW-KI-Werkstatt/DiffuMT

Journal refProceedings of Machine Learning Research (PMLR)2027

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑