发表机构
University of California, Santa Cruz; Johns Hopkins University(加州大学圣克鲁兹分校; 约翰霍普金斯大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文通过配对64 mT与3T扫描测量真实超低场MRI的信息上限,发现合成退化高估了可恢复信息,现有模型输出细节与受试者无关,PSNR/SSIM无法区分恢复与虚构,并发布测量协议以检验恢复能力。
AI 中文摘要
生成式超分辨率模型可以将便携式64 mT MRI图像转换为看起来像3T扫描的图像,该领域通常使用PSNR、SSIM和逐像素不确定性来评估这些模型,且大多数评估基于通过合成退化高场图像构建的配对数据。先前的工作承认这些模型会产生幻觉且该问题是不适定的,但据我们所知,尚无研究测量真实低场扫描实际包含的关于个体受试者的信息量。我们对此进行了测量。利用来自三个公共数据集的同一受试者的配对64 mT和3T扫描,以及一个在已知正确答案的测试上验证过的测量协议,我们发现,就全脑评估而言,真实64 mT扫描携带的个体特有结构仅在平面内约3至4毫米半间距以下,平面间则更粗糙。标准合成退化保留的受试者信息大约超出此上限1毫米,因此基于合成配对训练和基准测试的模型所评估的信息是真实扫描仪从未记录的。随后,我们在真实配对采集上测试了训练好的扩散模型和一个公开发布的外部模型;24次训练运行覆盖五种架构(GAN、扩散和Transformer系列)构成了审计范围。在每个可测量忠实度的受试者上,精细输出细节与受试者自身3T扫描的相关性并不高于与陌生人的相关性,而样本方差不确定性无法区分虚构结构与重建困难。由于PSNR和SSIM评分的是与参考图像的相似性而非细节是否属于受试者,因此以它们为指标的基准无法区分恢复与虚构。测量协议的代码将发布,以便对新模型的恢复能力声明进行测试。
英文摘要
Generative super-resolution models can turn portable 64 mT MRI into images that look like 3T scans, and the field evaluates them with PSNR, SSIM, and pixelwise uncertainty, most often on pairs built by synthetically degrading high-field images. Prior work acknowledges that these models hallucinate and that the problem is ill posed, but to our knowledge no study measures how much information about the individual subject the real low-field scan actually contains. We measure it. Using paired 64 mT and 3T scans of the same subjects from three public datasets, and a measurement protocol validated on tests whose correct answer is known in advance, we find that, judged over the whole brain, real 64 mT scans carry structure specific to the individual only down to approximately 3 to 4 mm half-pitch in plane, and coarser still through plane. Standard synthetic degradations preserve subject information roughly 1 mm beyond this ceiling, so models trained and benchmarked on synthetic pairs are evaluated on information that real scanners never record. We then test trained diffusion models and a publicly released external model on real paired acquisitions; 24 trained runs of five architectures (GAN, diffusion, and transformer families) give the coverage of the audit. On every subject where faithfulness can be measured, fine output detail is no more correlated with the subject's own 3T scan than with a stranger's, while sample-variance uncertainty does not distinguish fabricated structure from reconstruction difficulty. Because PSNR and SSIM score resemblance to a reference rather than whether detail belongs to the subject, a benchmark scored by them cannot tell recovery from fabrication. Code for the measurement protocol will be released so that recoverability claims can be tested for newer models.
Comments24 pages, 10 tables, 8 figures