arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

解剖结构忠实但时间上盲目:超声心动图左心室射血分数估计的归因审核

Ablation-Corrected Evaluation of Attribution Maps in Echocardiographic Ejection-Fraction Models

Hyunkyung Han, Min Jung Kim

arXiv 2607.13738首次发表:更新:

发表机构

School of Integrated Medicine, Yonsei University; Department of Radiology, Research Institute of Radiologic Science, Yonsei University College of Medicine(延世大学统合医学院; 延世大学医学院放射科学研究所放射科)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究超声心动图左心室射血分数估计模型的归因,微调VideoMAE Transformer和R(2+1)D CNN两个回归器,用多指标审核发现模型空间忠实但时间盲目,强调空间忠实不代表时间忠实,警示XAI验证并呼吁时间感知训练评估。

AI 中文摘要

背景与目标:深度视频模型从超声心动图估计左心室射血分数(EF),准确率接近专家水平,事后归因(Transformer的Chefer相关性、CNN的Grad-CAM)越来越多地用于验证模型“看对了地方”。但这些解释在空间和时间上是否忠实未经审核。由于EF由收缩末期(ES)和舒张末期(ED)帧定义,忠实的解释必须定位左心室(空间)和决定性帧(时间)。方法:我们在EchoNet-Dynamic上微调两个不同的EF回归器——一个自监督的VideoMAE Transformer和一个Kinetics预训练的R(2+1)D CNN,并沿三个轴用与架构匹配的归因进行审核:与左心室掩码的相关性交叉(IoR)、删除AUC以及ES/ED帧上的时间定位指数,每项相对于机会,在50项研究中每个案例有95%的置信区间。一个管段遮挡探针将归因失败与模型行为区分开来。结果:两个模型在解剖结构上都是忠实的——IoR分别比机会高2.91倍(VideoMAE)和1.98倍(R(2+1)D)——但在时间上是盲目:时间定位与机会无差异(0.97 - 1.00),并不比随机归因好。遮挡表明模型没有优先依赖ES/ED(0.90倍机会),所以时间盲目反映模型行为,而非归因假象。结论:空间忠实并不意味着时间忠实。归因可以证明解剖学基础,同时掩盖模型忽略临床决定性帧的事实——这对基于XAI的视频诊断模型验证是一个警示,并呼吁进行时间感知训练和评估。

英文摘要

Attribution maps for echocardiographic ejection-fraction models are evaluated by their overlap with an expert left-ventricular annotation, compared against a chance level that is computed from an area ratio rather than measured. We measure it. Two architectures trained on the same task attain overlap at 3.55 and 4.20 times measured chance, an eighteen percent difference a reader would take as the size of the gap between them. It is not. Replacing the annotated ventricle with a composition-matched surrogate changes one model's prediction six times more than replacing an equal-area control region and the other's twice, a factor of three; on sixty-two percent of cases for the second, neither the ventricle nor the attributed region moves the prediction at all. The reliability of this measurement, from two independent intervention batches, is 0.90 to 0.98. Overlap does not track the difference: scored against an ablation-derived label of which cases the model depends on the ventricle for, it reaches an area under the curve of 0.52 and 0.67. The design of the ablation also determines its answer, since ablating only the two annotated frames leaves the ventricle indistinguishable from a control region while ablating across all frames does not. A pediatric cohort reproduces all of these. We give the measured chance level, the symmetric-ablation protocol, the reliability estimate, and the ablation-prediction score as things to report alongside overlap.

Comments12 pages, 4 figures. v2: the chance level is now measured rather than computed, which corrects the ratios reported in v1. Adds a composition-matched ablation with an equal-area control, a reliability estimate for it, an ablation-derived score to report alongside overlap, seed repetition, and pediatric validation

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑