发表机构
Eurecat; University of Barcelona(欧洲研究与技术中心; 巴塞罗那大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文诊断扩散模型不确定性图失效的根源,发现是目标而非估计器所致,并提出无梯度的T-PT探针作为工具,在医学影像上以更低计算成本取得可比效果。
AI 中文摘要
扩散模型可以从基线预测后续医学扫描,但临床医生需要一张逐体素的地图,以了解该预测在哪些位置可以被信任。许多此类地图近似于Tweedie后验协方差的对角线,并针对其另一种近似进行评估,因此尚不清楚是估计器还是目标限制了它们。我们在十四个模型-语料库条件下的六个检查点上计算了精确对角线。Hutchinson在M=200时,其秩一致性至少为0.92,但在十四个条件中的四个条件下,精确对角线与去噪误差呈负相关,达到-0.13,因此一个忠实的估计器会重现这种反转。所有四个条件均为真实图像条件;在模型自身的样本上,这种反转并未出现,因此在生成样本上评估会美化这一系列方法。限制这些地图的是目标,而非估计器。随后,我们引入了Tweedie探针-切线(T-PT),这是一种无梯度的残差探针,它反复破坏一个模型支持的预测,并测量去噪器响应的逐体素方差。T-PT读取同一雅可比矩阵的不同泛函,其精确二阶形式在对角线反转之处与对角线排序一致;在三十次探针时,它返回的地图过于不稳定,无法重现该排序,而Hutchinson在M=5时已能重现,因此T-PT在此处并不构成对反转的反对证据。我们将其作为一种工具,而非更好的近似。在全分辨率脑部MRI上,我们测试的每个基于雅可比矩阵的估计器都会内存不足,T-PT在组织内的八个端点中的五个上领先于二十链蒙特卡洛集成,且在任何端点均未落后,网络评估次数减少了16倍;在整个体积上,集成领先,五十链缩小了组织差距。在肺部CT上,集成全程领先。两者在变化发生之处都失去了大部分区分能力,这一问题仍然悬而未决。
英文摘要
A diffusion model can predict a follow-up medical scan from a baseline, but a clinician needs a per-voxel map of where that prediction can be trusted. Many such maps approximate the diagonal of the Tweedie posterior covariance, and are evaluated against another approximation of it, so whether the estimator or the target limits them is unclear. We compute the exact diagonal on six checkpoints across fourteen model-corpus conditions. Hutchinson at M=200 tracks it at rank agreement of at least 0.92 everywhere, yet in four of the fourteen the exact diagonal is anti-correlated with the denoising error, reaching -0.13, so a faithful estimator reproduces that reversal. All four are real-image conditions; on the models' own samples the reversal does not appear, so evaluating on generated samples flatters this family. What limits these maps is the target, not the estimator. We then introduce Tweedie Probe-Tangent (T-PT), a gradient-free residual probe that corrupts one model-supported prediction repeatedly and measures the voxel-wise variance of the denoiser's response. T-PT reads a different functional of the same Jacobian, and its exact second-order form ranks with the diagonal wherever the diagonal reverses; at thirty probes it returns a map too unstable to reproduce that ranking, while Hutchinson at M=5 already reproduces it, so T-PT there is not evidence against the reversal. We offer it as an instrument, not a better approximation. On brain MRI at full resolution, where every Jacobian-based estimator we test runs out of memory, T-PT leads a twenty-chain Monte-Carlo ensemble on five of eight endpoints inside tissue and trails it on none, at 16x fewer network evaluations; over the whole volume the ensemble leads, and fifty chains close the tissue gap. On lung CT the ensemble is ahead throughout. Both lose most of their discrimination where the change is, which remains open.
Comments40 pages, 3 figures, 23 tables