医疗图像生成中的分词器-生成器耦合
Tokenizer-Generator Coupling in Medical Image Generation
- University College London(伦敦大学学院)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
该研究针对64×64低分辨率医疗图像,发现分词器与生成器的分离不合理,提出邻域条件预测增益指标,调优D3PM和SE-D3PM后FID-192大幅下降,验证了分词器-生成器耦合的重要性。
AI中文摘要:
潜在医疗图像生成器通常将分词器视为固定的预处理步骤。我们在64×64的ChestMNIST受控研究中测试这种分离是否合理,在共享潜在网格及连续潜在参考单元下,交叉离散分词器、生成器族和采样器设置进行实验。在该受控设置中,排名结果同时依赖于分词器、生成器和采样器:最优量化器会随生成器变化,基于验证集的采样器选择会改变生成器的表观排名。我们在三个随机种子下重新训练词汇量为1024的交互块,该交互块的有效性得以保留(9组成对量化器比较中有6组超过三个种子的标准差),并据此划定更宽的单种子网格。仅重建PSNR并非可靠的选择标准;我们提出一种无生成器的统计量——邻域条件预测增益,该指标可通过下游生成质量区分量化器族(秩AUC为1.00),而重建PSNR和边际令牌熵无法做到。在LFQ-1024上,对D3PM和SE-D3PM(基于保留验证集划分选择)进行调优后,它们在更低NFE下的默认FID-192从0.44/0.41降至0.09/0.10,且在不同种子间可复现;连续参考未进行等效采样器扫描。我们采用FID-192作为内部排名指标,其与标准FID-2048的斯皮尔曼相关系数为0.80,与无标签分类器双样本测试的相关系数为0.78。我们通过率-失真-可建模性框架解释结果,其中可建模性依赖于生成器、采样器和推理预算。所有实验均针对64×64的低分辨率医疗风格图像,采用无条件设置,使用非临床的基于FID的指标评估,且所有结论均限定于该场景。代码见此链接和此链接。
英文摘要:
Latent medical image generators usually treat the tokenizer as fixed preprocessing. We test whether this separation holds in a controlled ChestMNIST study at $64\times64$ that crosses discrete tokenizers, generator families, and sampler settings under a shared latent grid, with continuous-latent reference cells. The rankings depend jointly on the tokenizer, generator, and sampler. The best quantizer changes with the generator, and validation-based sampler selection changes the apparent generator ranking. We interpret this through a rate-distortion-modelability framing in which modelability is conditional on the generator, sampler, and inference budget. The interaction persists when the vocabulary-1024 block is retrained at three seeds, with 4 of 9 pairwise quantizer comparisons exceeding three seed standard deviations, including a reversal between LFQ and FSQ from MaskGIT to D3PM; the rest of the grid was trained at a single seed and is correspondingly less certain. Reconstruction PSNR alone is not a reliable selection criterion. On LFQ-1024, retuning the D3PM and SEDD samplers on a held-out validation split reduces FID-192 from 0.44/0.41 at the default budget to 0.09/0.10 at lower NFE, replicated across seeds, although the continuous references were not given an equivalent sampler sweep. FID-192, our internal ranking metric, ranks consistently with standard FID-2048 (Spearman 0.89) and with a label-free classifier two-sample test (0.86). All experiments are unconditional, use low-resolution $64\times64$ medical-style images and are evaluated with non-clinical FID-based metrics, and our claims are limited to that setting. Implementations are released at https://github.com/liamchalcroft/medtokenizers and https://github.com/liamchalcroft/medlatents.