超声心动图视频分割中时间可解释性的定量评估框架
A Quantitative Evaluation Framework for Temporal Explainability in Echocardiographic Video Segmentation
- University of Toronto(多伦多大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文提出一个定量评估框架,用四个指标衡量超声心动图视频分割中Grad-CAM时间可解释性,发现中间层解释不稳定,时间瓶颈更稳定,并推动时间感知XAI方法。
AI中文摘要:
深度学习在超声心动图视频分割中已达到最先进的性能,越来越多的模型融入了时间信息。然而,时间可解释性的定量评估在很大程度上仍未得到探索。我们提出了一个定量框架,使用四个互补指标来评估Grad-CAM解释,这些指标分别衡量时间一致性、显著性运动、解剖重叠和时间重叠。利用EchoNet-Dynamic,我们比较了基线2D U-Net与在多个时间步长上训练的ConvLSTM U-Net模型。虽然所有模型的分割性能保持相当,但中间层ConvLSTM解释表现出显著较低的显著性一致性和更大的质心运动,相较于最终预测解释。在所有步长中,时间瓶颈(Temporal Bottleneck)解释显著比编码器瓶颈(Encoder Bottleneck)解释更稳定,而最终ConvLSTM解码器3(Decoder3)解释与2D U-Net的解释大致相当。重要的是,传统的逐帧解释指标无法确定中间层解释的变化是反映了有意义的时间特征演化还是解释的不稳定性。这些发现为时间可解释性建立了一个初步的定量框架,并推动了在医学视频模型中明确考虑演化表征的时间感知XAI方法的发展。
英文摘要:
Deep learning has achieved state-of-the-art performance in echocardiographic video segmentation, with an increasing number of models incorporating temporal information. However, quantitative evaluation of temporal explainability remains largely unexplored. We propose a quantitative framework for evaluating Grad-CAM explanations using four complementary metrics measuring temporal consistency, saliency motion, anatomical overlap, and temporal overlap. Using EchoNet-Dynamic, we compare a baseline 2D U-Net with ConvLSTM U-Net models trained across multiple temporal strides. While segmentation performance remained comparable across all models, intermediate ConvLSTM explanations exhibited substantially lower saliency consistency and greater centroid motion than final prediction explanations. Temporal Bottleneck explanations were significantly more stable than Encoder Bottleneck explanations across all strides, while final ConvLSTM Decoder3 explanations were broadly comparable to those of the 2D U-Net. Importantly, conventional frame-wise explanation metrics cannot determine whether variation in intermediate explanations reflects meaningful temporal feature evolution or explanation instability. These findings establish a preliminary quantitative framework for temporal explainability and motivate temporal-aware XAI methods that explicitly account for evolving representations in medical video models.