发表机构
The University of Texas at Austin(德克萨斯大学奥斯汀分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究构建了最优传输修正流系统,在威斯康星乳腺癌数据集上实现了无监督恶性评分与归因,发现胸部X光的无标签最优传输热图无法真实定位疾病,仅监督式Grad-CAM表现良好,并贡献了相关方案与基准。
AI 中文摘要
生成模型为可解释临床AI提供了一条途径:无需探测分类器,而是对健康和患病患者的分布进行建模,并从二者之间的几何关系中提取解释。我们构建了这样一个系统——在两种临床分布之间训练的最优传输修正流,并以此提出该领域极少检验的关键问题:生成的解释热图是否真的能定位疾病?在表格型肿瘤生物标志物(威斯康星乳腺癌数据集)上,单一的流可生成每位患者的反事实样本、无监督恶性程度评分(AUROC为0.91;5次随机种子下的平均值为0.93±0.01),以及与有监督分类器一致的无标签归因(相关系数r≈0.5)——这是一种紧凑、可靠的可解释性引擎,尽管其预测能力从未超过逻辑回归。在胸部X光片上,我们发现传输热图是一种群体层面的信号,而非定位器;基于重建的保留身份变体确实能定位合成病变(指向游戏得分为0.52),但在真实RSNA放射科医师标注框上,其表现退化为随机水平,仅监督式Grad-CAM保持在随机水平之上。核心结果是存在合成到真实的差距:在植入病变上看似有说服力的无标签热图,并非真实定位的证据。我们贡献了可复用的生成解释最优传输方案,以及用于压力测试其是否能定位的受控基准。
英文摘要
Generative models promise a route to explainable clinical AI: rather than probe a classifier, model the distributions of healthy and diseased patients and read explanations off the geometry between them. We build such a system - an optimal-transport rectified flow trained between two clinical distributions - and use it to ask a pointed question the field too rarely tests: do the resulting explanation heatmaps actually localize disease? On tabular tumour biomarkers (Breast Cancer Wisconsin) a single flow yields per-patient counterfactuals, an unsupervised malignancy score (AUROC 0.91; 0.93 +/- 0.01 across five seeds), and a label-free attribution that agrees with a supervised classifier (r ~ 0.5) - a compact, honest interpretability engine, though it never out-predicts logistic regression. Moving to chest X-rays, we show the transport heatmap is a population-level signal, not a localiser; a reconstruction-based, identity-preserving variant does localize synthetic lesions (pointing game 0.52), yet on real RSNA radiologist boxes it collapses to chance while only supervised Grad-CAM stays above it. The central result is a synthetic-to-real gap: label-free heatmaps that look compelling on planted lesions are not evidence of real localisation. We contribute a reusable optimal-transport recipe for generative explanations and a controlled benchmark for stress-testing whether they localize.