什么造就了好的医学图像分词器?重新思考医学图像分词中的重建与生成
What Makes a Good Medical Image Tokenizer? Rethinking Reconstruction and Generation in Medical Image Tokenization
浏览论文内容
中文总结 AI 辅助
该研究系统评估了医学图像分词器,发现重建与生成性能强相关,现代分词器未充分利用潜在空间,离散量化可保留下游分类能力。
中文摘要 AI 辅助
潜在扩散模型如今主导着医学图像生成,而每一个此类流程都依赖于一个将图像压缩为潜在编码以供图像生成操作的分词器。因此,分词器的选择制约着从重建保真度、生成质量到可供下游分析使用的表示等所有下游任务。然而,医学图像流程通常采用自然图像中的分词器,其假设是这些分词器的行为可以迁移。但这一假设在医学图像领域从未被检验,在该领域中数据集规模要小几个数量级,且图像间的样本方差要低得多。我们对医学图像分词器进行了系统评估,在三个压缩因子下,跨越十个模型家族、十二个数据集,评估了三十种配置,涵盖重建、生成、潜在几何、下游分类和记忆化。我们发现:(1) 图像重建和生成的性能强相关,这与先前关于自然图像的报告不同;(2) 现代分词器几乎使用了其所有码本条目,但仍留下大部分潜在空间未使用;(3) 训练集记忆化程度较轻,且更强的潜在空间压缩可进一步抑制;(4) 离散量化能在很大程度上保留下游分类能力,而无查找方案是主要例外。
英文摘要
Latent diffusion models now dominate medical image generation, and every such pipeline rests on a \emph{tokenizer} that compresses images into the latent codes for image generation to operate on. Thereby, the tokenizer choice bounds every downstream task from reconstruction fidelity and generation quality to the representations available for downstream analysis. Yet, medical imaging pipelines routinely utilize tokenizers from natural imaging on the hypothesis that their behavior carries over. However, this is an assumption never tested in the medical imaging regime, where datasets are orders of magnitude smaller and images exhibit far lower inter-sample variance. We present a systematic evaluation of medical image tokenizers evaluating thirty configurations across ten model families on twelve datasets at three compression factors, spanning reconstruction, generation, latent geometry, downstream classification, and memorization. We find that (1) performance on image reconstruction and generation strongly correlate, unlike prior reports on natural images; (2) modern tokenizers use nearly all of their codebook entries, but still leave most of the latent space unused; (3) training-set memorization is mild and is further suppressed by stronger latent space compression; and (4) discrete quantization can largely preserve downstream classification, with lookup-free schemes being the main exception.
发表机构
- Technical University of Munich(慕尼黑工业大学)
- Munich Center for Machine Learning, Technical University of Munich(慕尼黑工业大学慕尼黑机器学习中心)
- Klinikum rechts der Isar, Technical University of Munich(慕尼黑工业大学伊萨尔河右岸医院)
- ELLIS Institute Finland(芬兰ELLIS研究所)
- Aalto University(阿尔托大学)
- Imperial College London(伦敦帝国理工学院)
机构由 AI 辅助整理,请以论文原文为准。